Skip to main content
Back to timeline
Scientific ReportsSource publication:

LLM pairwise comparison of historical proposals at three Spallation Neutron Source beamlines ranks them in positive correlation with human ranking, matches human reviewers at flagging high-publication-potential proposals, and costs over two orders of magnitude less

Synopsis

Using historical general user proposals from three beamlines at Oak Ridge National Laboratory's Spallation Neutron Source (EQ-SANS, CNCS, POWGEN), the study has large language models judge every pair of proposals within a run cycle, converts the win-lose outcomes into rankings with the Bradley-Terry model, and finds that LLM rankings correlate positively with human rankings (Spearman ρ about 0.2-0.8, rising to ≥0.5 after 10% outlier removal), show no statistically significant difference from human review in identifying proposals with high publication potential, cost over two orders of magnitude less, and support linear-complexity proposal similarity analysis via embedding models.

AI-generated editorial illustration: LLMs can assist with proposal selection at large user facilities

Interpretation

It introduces and validates a pipeline in which an LLM performs pairwise preference comparisons and the Bradley-Terry model converts win-lose outcomes into proposal rankings; on historical proposals from three SNS beamlines over the past 20 run cycles, LLM rankings correlate positively with human rankings, with Spearman ρ about 0.2-0.8, rising to ≥0.5 after removing 10% outliers. Prior LLM work on peer review mostly assessed the absolute quality of individual manuscripts, whereas proposal selection cares about relative strength among proposals; this work decomposes the task into pairwise preference judgments and estimates relative strength from win-lose outcomes via Bradley-Terry, sidestepping the weak inter-proposal correlation of human individual scoring. Based on historical proposals and human scoring records from three representative SNS beamlines (EQ-SANS, CNCS, POWGEN) over the past 20 run cycles, with per-cycle Spearman rank correlations and reported changes after outlier removal.

Using publication records associated with proposals as a reference, LLM and human rankings show no statistically significant difference in identifying proposals with high publication potential: for EQ-SANS the publication metric is 0.481±0.079 (LLM) versus 0.474±0.110 (human), for POWGEN 0.542±0.106 versus 0.514±0.093, and for CNCS both are 0.516. The work ties ranking effectiveness to the subsequent publication output of proposals, constructing a publication metric from discounted publication counts (a paper linked to K proposals counts as 1/K), providing a reference for ranking quality independent of human scores. Only run cycles with at least 4 proposals having associated publications are included to improve statistical significance, and means and standard deviations of the publication metric are reported per beamline; the authors also note that historical acceptance decisions were based on human ranking, so results are expected to be biased toward human ranking.

Cost estimates based on U.S. Bureau of Labor Statistics wage data and model token prices put human review at about $54.9 per proposal review versus about $0.0046 per LLM pairwise preference comparison; for a typical proposal pool of N∈[30,70], human cost is 346 to 823 times the LLM cost, i.e., the LLM approach costs about 0.12%-0.29% of human reviewers using individual scoring. Prior discussion of LLM-assisted review largely stayed at the capability level; this work places the O(N²) workload of pairwise preference and the O(N) workload of human individual scoring in one cost framework and quantifies the gap. Cost inputs come from U.S. Bureau of Labor Statistics 2025 data (postsecondary engineering teachers at $119,340/year, 41.8 average weekly working hours) and the token prices of the model used (Gemini-2.5-flash), with token usage statistics from all pairwise comparisons run in this work; the authors note that the roughly 1 hour per review is a rough estimate, but the cost gap is large enough in magnitude.

With an LLM embedding model (Qwen3-embedding-8b, 4096 dimensions) turning each proposal into a vector, proposal similarity can be analyzed via cosine similarity at O(N) complexity rather than the O(N²) human effort; in EQ-SANS cycles 25A and 25B, the highest-similarity points correspond to a revised resubmission and to two proposals on the same topic from different principal investigators. Similarity checks between proposals (duplicate or overlap detection) are hard for human reviewers to perform consistently due to cognitive load and fatigue; this work turns the task into a single forward pass plus batch dot products, making large-scale similarity analysis routine. Similarity-matrix heatmaps are shown for proposals from EQ-SANS cycles 25A and 25B, with the highest-similarity points identified as a resubmission and a same-topic pair from different principal investigators; the similarity threshold must be set manually before human verification.

Perspective

The results target general user proposal selection at large user facilities such as SNS, using historical proposals and human scoring records within the same run cycle; the pipeline runs PDF through OCR to markdown, has an LLM make pairwise preference judgments, estimates relative strength with the Bradley-Terry model, and optionally adds embedding-based similarity analysis. The authors state the framework is not limited to SNS and can be applied to other ORNL facilities, other national laboratories, and funding agencies such as DOE, NSF, and NIH; for very large proposal batches, the comparison matrix can be sparse and the choice of paired proposals optimized. Proposed next steps include scaling experiments to more facilities, testing different LLMs, and fine-tuning a model on past proposal and publication data to predict publication likelihood.

The publication metric covers only accepted proposals, and historical acceptance decisions were based on human ranking, so the metric is expected to be biased toward human ranking; the authors also note the lack of ground truth for proposal ranking. In the cost estimate, roughly 1 hour per review is a rough estimate that the authors say carries most of the uncertainty. In similarity analysis, the threshold must be set manually, and the resubmission and same-topic cases come from specific cycle examples. In addition, this is a full-text parse, and the specific numerical details in figures (such as point-by-point values in per-cycle scatter plots and heatmaps) are not given individually in the body text, so per-cycle verification still requires consulting the original figures.

Sources