Skip to main content
Back to timeline
arXivSource publication:

RECLAIM benchmark: four agents reproducing results on 100 NeurIPS 2025 papers succeed on only 41% of Run-tier, 27% of Retrain-tier, and 15% of Reimplement-tier papers at best

Synopsis

The authors introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences, fixing in advance for each paper the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget, and dividing difficulty by what authors released (Run tier has code, data, and weights; Retrain tier lacks weights so the agent trains the model; Reimplement tier lacks code so the agent writes it), with a separate language model grading runs from logs and outputs rather than agents' reports; four agents run once per paper, and the best agent in each tier reproduces only 41%, 27%, and 15% of papers respectively, failed attempts use on average 29% of their budget, and the most common error is writing the method without checking any part against the pape

Source-provided article image: RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?
Figure 1 ·

Figure 1: Overview of RECLAIM and its evaluation pipeline. The figure shows an example run with Qwen3.6-27B on the TTPL paper ( Chen et al., 2025d ) .

arXiv

Interpretation

RECLAIM turns ML paper reproduction into a benchmark that is defined in advance and rebuildable yearly: for 100 NeurIPS 2025 papers it fixes the result to reproduce, the success criterion, and a GPU-hour budget. Prior evaluations of agent research ability often rely on task setups or post hoc judgment; here the reproduction target, success criterion, and resource cap are fixed before running, and the benchmark can be rebuilt from new conferences. Based on a benchmark of 100 NeurIPS 2025 papers; the abstract states the scale, source conference, and the three pre-fixed elements.

Difficulty is tiered by what authors released: Run tier includes code, data, and weights; Retrain tier lacks weights so the agent trains the model; Reimplement tier lacks code so the agent writes it. It maps 'what the authors released' directly onto task difficulty, so reproduction ability can be measured separately under different levels of release completeness. The abstract explicitly defines the three tiers and the corresponding differences in released artifacts.

Grading is done by a separate language model from logs and outputs rather than by trusting agents' reports. It shifts the basis of judgment from agents' self-reporting to checkable run logs and outputs, reducing self-report bias in the results. The abstract states the grading agent and its basis, but does not give the grading model's identity or agreement data with human judgment.

Four agents run once per paper; the best per tier reproduces 41% of Run-tier, 27% of Retrain-tier, and 15% of Reimplement-tier papers; failed attempts use on average 29% of their budget, and the most common error is writing the method without checking any part against the paper's numbers, in 63 of 400 runs. It provides tier-wise reproduction success rates under a budget constraint and a typical failure mode, showing success falls as less is released and that most failures do not exhaust the budget. The abstract reports success rates for four agents running once per paper, the average budget fraction used, and an error count of 63 out of 400 runs.

Perspective

This work is aimed at researchers and engineering teams who need to evaluate or build research agents, and it applies to settings where a paper and its released artifacts are the input and a specified result must be reproduced within a fixed GPU-hour budget. The tiered design lets one benchmark characterize reproduction ability separately under 'code, data, and weights available', 'weights missing', and 'code missing' conditions, and it can be rebuilt yearly from new conferences to track progress over time.

The abstract does not state the identity of the language model used for grading, its agreement with human judgment, or variation across repeated runs, so the stability of the 41%, 27%, and 15% figures is unclear. Failed attempts use on average only 29% of their budget, indicating most failures are not budget exhaustion, but the abstract does not explain why agents stop early. In addition, this summary is based only on the abstract, without the paper's body, figures, or appendix, so the full benchmark construction details and error taxonomy still need to be checked in the original.

Sources