CroissantMiner releases a 602-paper benchmark: single-pass full-context extraction beats four agentic architectures, with the best system reaching only 0.709 composite
Related research and updatesSynopsis
The work builds CroissantMiner, a benchmark pairing 602 ML dataset papers (102 with human-validated gold annotations, 500 with LLM-generated silver annotations) with the full 30-field Croissant 1.1 schema, and evaluates 24 extraction systems under a two-tier framework combining rule-based scoring with a human-audited LLM judge, finding that single-pass full-context extraction outperforms all four evaluated agentic architectures on every backbone and that long-form RAI fields are hardest, with the strongest system reaching only 0.709 composite.
Figure 1: The CroissantMiner data-creation pipeline. (1) Corpus: ML dataset papers from top-downloaded Hugging Face datasets (vision, NLP, audio, etc.). (2) Extraction: a single-pass LLM generates 30-field Croissant metadata per paper, forming the silver split and the pre-fills for the gold split. (3) Annotation: annotators rate each pre-fill against the source paper on a 3-level rubric (Correct / Partially Correct / Not Correct) with failure-mode labels (e.g., hallucination, incomplete), producing 9,595 ratings over 3,060 cells. (4) Adjudication: majority vote resolves 95% of cells; senior-author review resolves the rest.
arXivInterpretation
The first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema: 602 dataset papers paired with the 30-field Croissant 1.1 schema, of which 102 carry gold annotations from 22 annotators providing 9,595 ratings and 500 carry LLM-generated silver annotations. Prior evaluation data was scarce and schema-specific (e.g., Masader, CardBench) and did not produce Croissant-conformant output; this benchmark is the first to pair scientific papers with Croissant-RAI-conformant gold annotations. Gold labels use a pre-fill plus human-verification protocol with at least three independent ratings per cell, majority voting resolving 95.0% of cells and senior adjudication for the rest; 83.4% of ratings judged the pre-filled value fully correct and 12.1% supplied an explicit correction.
With the backbone held fixed, all four agentic architectures score below single-pass extraction: on Claude Sonnet 4.6, ReAct reaches 0.652, Parallel Specialists 0.647, Triage + Critique 0.624, and Locator-Extractor 0.566, all below the 0.709 of single-pass, with differences significant under the paired bootstrap and Wilcoxon signed-rank test after BH-FDR correction. The result covers four mainstream decomposition patterns (role-based specialists, plan-then-execute with self-critique, retrieval-based context narrowing, and dynamic tool-use loops) and repeats on GPT-5.4 and Gemini 3.1 Pro backbones, attributing the gap to the decomposition strategy rather than a single design choice. Evaluated on an 88-paper held-out test split with 2,000-replicate paper-clustered bootstrap confidence intervals; on RAI cells where the paper documents a value, single-pass returns an empty value for only 2.4% of cells while the four agentic variants miss between 10.9% and 14.8%.
Long-form RAI fields are the hardest extraction targets because they require synthesizing information across sections rather than copying from one location, and agentic architectures degrade more on RAI than on core fields. The analysis locates difficulty in field type and documentation sparsity rather than paper length: single-pass scores fall with how sparsely a paper is documented, while page count has no measurable effect. In the gold split 27.7% of cells have no documented value, nearly three times higher for RAI fields (35.6%) than core fields (11.9%); analysis of 117 sampled cells where agentic systems scored below single-pass shows leaving documented fields empty at 44%, incomplete answers at 34%, and judge noise at only 5%.
Single-pass extraction is also the most cost-efficient variant: in matched-backbone comparisons, agentic configurations cost 1.2 to 7.8 times as much per paper as single-pass without improving composite scores. The joint cost and accuracy result indicates that task decomposition did not buy quality gains across the budget tiers evaluated. Costs are estimated from recorded token usage at standard public list prices, applying the same pricing basis to single-pass and agentic runs and excluding papers with unrecorded token usage.
Perspective
The benchmark targets English-language ML dataset papers and the 30-field Croissant 1.1 schema, and is meant for an author-in-the-loop draft-then-review workflow: the system fills the 30 fields from the manuscript and flags low-confidence fields back to authors, who review the fields where the model is empirically weakest (e.g., sc:publisher, rai:dataCollectionType, rai:dataAnnotationAnalysis). A composite of 0.709 is not enough for unattended metadata production but is a useful starting point for an author-facing verification tool; centralized deployment at journals and community repositories, combined with deterministic tooling for structurally extractable fields, amortizes the per-paper cost across many users. The silver split supports distillation, fine-tuning, and corpus-level coverage-trend studies, while the gold split supports per-field accuracy estimates.
Gold labels began as Claude Sonnet 4.5 pre-fills; although human verification and an author audit revised 191 of 3,060 cells, the residual structural pattern of what the seed model marks as filled versus null is inherited by the gold, which is why that model is excluded from the ranking. Re-annotating 10 papers from GPT-5.4 pre-fills left other systems' ranking stable, but the seed model and, to a lesser degree, its family gain, so cross-family comparisons call for care. Tier-2 scoring relies on a single LLM judge, GLM-5, which matches human consensus in 71.5% of cells, close to the 72.0% agreement between two human raters, and is usually more lenient when it disagrees, so judge-based scores do not replace full human evaluation. Several RAI fields have very few documented cells in the test split (as few as two for rai:dataImputationProtocol), so results for these fields should not be over-read, and pretraining contamination cannot be ruled out. The benchmark covers English-language papers and Croissant 1.1; extension to other languages, domains, and future schema versions is left to future work, and vision-language preprocessing that reads tables and figures as images remains untested.
