Skip to main content
Back to timeline
arXivSource publication:

184 authors labeled what inspired 207 papers, and the best retriever recovers only 48% of those catalyst papers from a 191K corpus

Synopsis

The authors built ScholarCatalyst, a benchmark in which 184 researchers annotated 894 early-stage research questions and the prior "catalyst" papers that did or could have advanced 207 recent computer science projects, then evaluated retrieval over a 191K-paper corpus restricted to work published before each project; the strongest system reached only 0.48 Recall@20, agentic search (0.42) did not beat embedding retrieval, and even Claude Fable 5.1, whose training data may include the source papers, reached only 0.51.

AI-generated editorial illustration: ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Interpretation

The paper introduces and releases ScholarCatalyst, the first scientific literature retrieval benchmark grounded in authors' firsthand accounts of which prior papers inspired their projects. Existing scientific retrieval benchmarks (SciFact, DORIS-MAE, ScholarQABench, LitSearch, MIR) evaluate claim support, facet-level relevance, or citation intent, whereas this work elicits early-stage research questions and inspiration judgments that the published record largely omits, including papers the authors had not encountered at the time. 184 researchers, 207 computer science and AI papers from 2025-2026, 894 queries (207 core research queries and 687 subfield-specific queries), a 191K-paper corpus, and author-provided positives, hard negatives, and rationales; authors rated the reconstructed questions accurate or mostly accurate for 98.1% of CoreQ and 95.9% of SubQ instances.

Current retrieval systems miss most author-credited inspiration papers, and agentic search does not improve on embedding retrieval. The work is the first to compare sparse, dense, and multi-vector retrievers against three search-agent designs on the same inspiration-retrieval task, and it attributes the agent gap to candidate coverage rather than reasoning. General-purpose dense retrievers perform best yet place 39% of CoreQ and 51% of SubQ gold papers in the top 20; scientific-document retrievers (SPECTER2, OpenScholar) trail by at least 16 points on CoreQ and 27 points on SubQ at Recall@20; BM25 and LateOn trail by 15-16 and 18-25 points; the tool-calling agent reaches 0.42 versus 0.48 for the retriever alone, and the grep agent finds only 8% of gold query-paper pairs during search versus 46% for the tool-calling agent.

Similarity and citation records are unreliable signals of inspiration, and inspiration relations are diverse and context-dependent. A nine-type taxonomy of author rationales shows inspiration often comes from an adapted method, supporting evidence, generalization to a new setting, or a motivating limitation, with more than half of rationales carrying multiple types; hard negatives are at least as similar to the query as positives on both lexical and semantic measures. 663 key-inspiration rationales received 1.69 labels on average, with 55.8% receiving two or more; 43.6% of SubQ positives are not cited by the source paper and 67.8% of those share no references with it; only 46.2% of CoreQ and 34.0% of SubQ positives carry a uses or extends citation-intent label; one prior paper served as a naive baseline, a cross-modality generalization target, and a structural premise across three unrelated source papers.

The bottleneck is knowing what to search for rather than understanding the papers already found. Candidate injection, context ablations, and query-rewriting experiments separate retrieval coverage from recognition performance. The retriever's top 100 covers 70% of gold papers and agents recover almost none of the rest, with about half of gold papers returned by no search during an agent's run; adding the source title and abstract raises CoreQ Recall@20 by 0.11 while adding the source bibliography raises it by 0.35 (the bibliography alone contains 95% of CoreQ gold papers); letting an agent read beyond abstracts changes Recall@20 by at most 0.04; single-query expansion and multi-query generation yield no meaningful gain, and both HyDE variants substantially reduce recall.

Perspective

The benchmark targets computer science and AI papers from 2025-2026, uses a 191K-paper corpus with a per-query temporal cutoff at the source paper's publication date, and evaluates systems with no web access and with model knowledge cutoffs preceding all source papers; it is meant to measure whether a system can recover author-credited inspiration papers for an early-stage research question, and its automated pipeline lets it grow with newly published work. The authors name two milestones ahead: expert-level retrieval within a field, and retrieval beyond any individual expert across many fields, since 49.5% of the papers their authors credit come from outside the project's subject area.

Labels are authors' retrospective judgments about completed projects and are subject to hindsight; the performance ceiling of the task remains an open question, with a preliminary comparison on one paper showing three coauthor annotators marking about seven positives per thread recovered 43-60% of the first author's labels at 84-88% precision. The 191K-paper corpus is far smaller than arXiv's three million articles or Semantic Scholar's 225M papers, so retrieval here is an easier version of the real problem. In addition, Claude Fable 5.1's training data may include the source papers, so its 0.51 R@20 serves only as a lenient reference rather than a comparable result.

Sources