Public articles linked to the same research event.
arXiv The work builds CroissantMiner, a benchmark pairing 602 ML dataset papers (102 with human-validated gold annotations, 500 with LLM-generated silver annotations) with the full 30-field Croissant 1.1 schema, and evaluates 24 extraction systems under a two-tier framework combining rule-based scoring with a human-audited LLM judge, finding that single-pass full-context extraction outperforms all four evaluated agentic architectures on every backbone and that long-form RAI fields are hardest, with the strongest system reaching only 0.709 composite.
The work builds CroissantMiner, a benchmark pairing 602 ML dataset papers (102 with human-validated gold annotations, 500 with LLM-generated silver annotations) with the full 30-field Croissant 1.1 schema, and evaluates 24 extraction systems under a two-tier framework combining rule-based scoring with a human-audited LLM judge, finding that single-pass full-context extraction outperforms all four evaluated agentic architectures on every backbone and that long-form RAI fields are hardest, with the strongest system reaching only 0.709 composite.
The work builds CroissantMiner, a benchmark pairing 602 ML dataset papers (102 with human-validated gold annotations, 500 with LLM-generated silver annotations) with the full 30-field Croissant 1.1 schema, and evaluates 24 extraction systems under a two-tier framework combining rule-based scoring with a human-audited LLM judge, finding that single-pass full-context extraction outperforms all four evaluated agentic architectures on every backbone and that long-form RAI fields are hardest, with the strongest system reaching only 0.709 composite.
The work builds CroissantMiner, a benchmark pairing 602 ML dataset papers (102 with human-validated gold annotations, 500 with LLM-generated silver annotations) with the full 30-field Croissant 1.1 schema, and evaluates 24 extraction systems under a two-tier framework combining rule-based scoring with a human-audited LLM judge, finding that single-pass full-context extraction outperforms all four evaluated agentic architectures on every backbone and that long-form RAI fields are hardest, with the strongest system reaching only 0.709 composite.