SciMuse generates personalized research ideas from a 58-million-paper knowledge graph and an LLM; over 100 research group leaders rated 4,400+ ideas at a mean of 2.40 on a 5-point scale, with 24.9% rated 4 or 5
Synopsis
The authors introduce SciMuse, which generates personalized research ideas using a knowledge graph of 58 million papers and a large language model, and they conducted a large-scale evaluation in which more than 100 research group leaders spanning the natural sciences to the humanities rated over 4,400 personalized ideas by level of interest; expert ratings were modest (mean 2.40 on a 5-point scale, most common rating 1) while 24.9% of ideas were rated 4 or 5, supplying knowledge-graph-selected concept pairs did not improve expert-rated interest over a titles-only GPT baseline, high-citation-predicted pairs even showed a weak tendency (1.
Figure 1: Overview of the four main components of SciMuse : an evolving impact-augmented knowledge graph, a method for generating personalized research suggestions, a large-scale expert evaluation across diverse scientific fields, and methods for predicting the scientific interest of research suggestions.
arXivInterpretation
The paper introduces SciMuse, a system that generates personalized research ideas by combining a knowledge graph of 58 million papers with a large language model. Relative to approaches relying only on an LLM or literature search, this work plugs a large-scale knowledge graph in as the source of concepts and explicitly targets personalization. The abstract states the knowledge graph scale (58 million papers) and the components used (knowledge graph plus LLM), which is system-description-level evidence.
More than 100 research group leaders from the natural sciences to the humanities rated over 4,400 personalized ideas by interest, with modest overall ratings: mean 2.40 on a 5-point scale, most common rating 1, while 24.9% of ideas were rated 4 or 5. The work provides a large-scale, interdisciplinary expert-rating dataset collected from research group leaders, turning the question of whether AI-generated ideas are interesting from speculation into quantified expert judgment. Evidence comes from over 100 experts rating more than 4,400 ideas, with sample scale and rater identity stated in the abstract, constituting large-scale expert-evaluation evidence.
Concept pairs selected using the knowledge graph did not improve expert-rated interest over a titles-only GPT baseline, and high-citation-predicted pairs even showed a weak tendency (1.94σ) toward lower interest than random pairs. This controlled result directly tests the assumption that graph-selected concepts improve idea quality and reports a directional signal contrary to intuition. Evidence comes from comparisons against a titles-only GPT baseline and random concept pairs, with the high-citation-predicted difference reported as a weak 1.94σ tendency.
Graph features can be used to control properties of ideas, and idea interest can be predicted with both a supervised neural network based on graph features and a zero-shot ranking approach based on an LLM. Even though graph concepts did not raise generation quality, the work redirects graph features toward controllability and interest prediction, training and validating predictors on the expert-rating dataset. Evidence comes from two prediction approaches (a supervised neural network and zero-shot LLM ranking) evaluated on the same expert-rating dataset; the abstract does not report specific prediction accuracy values.
Perspective
The results are aimed at researchers and tool builders who want AI assistance in choosing research topics: SciMuse's generation pipeline, graph-feature control, and interest-prediction methods can be reused where a comparable paper knowledge graph and expert-rating data exist. The evaluation setting is research group leaders rating personalized ideas by interest, so the conclusions apply to the metric of expert subjective interest, not to feasibility, correctness, or eventual research output.
The abstract does not report specific prediction accuracy, the comparison between the supervised neural network and zero-shot LLM ranking, or the statistical details of the 1.94σ weak tendency; moreover, the coexistence of a mean rating of 2.40 with 24.9% high-rated ideas suggests idea quality may be highly uneven, so readers judging the method's merits still need the figures and statistical tests in the full text.
