AutoDataBench scores Kimi-3 highest at 60.68 overall, yet the retrieval leader on in-distribution data ranks last out of distribution
Synopsis
The work introduces AutoDataBench, a controlled testbed that isolates data intelligence under fixed training pipelines and resource budgets across tool-use data diagnosis and repair, retrieval data organization, and knowledge injection, evaluates seven frontier LLMs on iterative data decisions, finds that agent-curated data approaches expert references on tool use and retrieval while strong in-distribution scores do not guarantee out-of-distribution transfer, and shows that reusing optimization trajectories as mid-training data improves downstream coding performance.
Interpretation
AutoDataBench isolates data decisions from training frameworks, hyperparameters, and compute by fixing pipelines across three tasks: tool use trains Qwen2-1.5B-Instruct on roughly 40k examples corrupted by seven probabilistic operators, retrieval trains a 6-layer MiniLM with InfoNCE contrastive learning on the RLHN collection, and knowledge injection distills a frozen teacher into 13B Talkie with offline privileged-context distillation. Existing auto-research benchmarks often entangle multiple routes to improvement, making it hard to attribute which research capability drives a result; this work fixes training pipelines, tools, and source datasets so agents only change data, and adds hidden out-of-distribution evaluation to identify benchmaxxing. Each task specifies source data scale, base model, training objective, and resource caps (one A800 GPU, 24 hours, at most 20 training-and-evaluation rounds), with both random baselines and expert references.
Seven frontier models approach expert-level tool-use data repair: mean in-distribution accuracy is 80.82–81.98, above the random baseline of 70.45 and close to the oracle-clean reference of 81.65; on BFCL all exceed the random baseline of 55.18 but remain below the expert reference of 60.23. This indicates agents can discover and repair structural and semantic errors in arguments, function labels, schemas, and query-call alignment without being told the corruption annotations, while a transfer gap remains. Each model-task combination has three independent standard runs, 63 runs in total, with three-run means and sample standard deviations plus separately reported out-of-distribution scores.
Retrieval shows a clear in-distribution versus out-of-distribution rank reversal: GPT-5.6-Sol reaches a mean in-distribution score of 40.46, close to the expert reference of 40.62, but scores lowest out of distribution at 26.15; Qwen-3.7-Max leads out of distribution at 30.18, while all models fall below the expert's 34.76. This directly demonstrates that data strategies favored by target-task feedback need not retain their advantage on unseen distributions, providing measurable evidence of benchmaxxing. Out-of-distribution scores are withheld throughout optimization and measured only at checkpoints selected by in-distribution feedback, averaged over five held-out retrieval datasets.
Reusing AutoDataBench optimization trajectories as mid-training data, adding 8M trajectory tokens (upsampled tenfold to 1.6% of a 5B-token budget), improves all five coding evaluations, with CRUXEval input prediction rising from 72.00 to 74.12 and SWE-bench Multilingual from 27.67 to 33.33, while MBPP and LiveCodeBench gains are smaller. The benchmark is not only an evaluation tool but also a source of training data, and the gains concentrate on tasks requiring backward reasoning from program behavior and diagnosis of unfamiliar repositories, which resemble the exploratory structure of the trajectories. Control and treatment arms share the same 5B-token mixture, identical SFT data, and the same random seed, differing only in whether auto-research trajectories are present during mid-training.
Perspective
The testbed targets researchers and engineering teams who want to examine data decisions in isolation under fixed training pipelines, and it applies to tool-use, retrieval-embedding, and knowledge-injection data problems under a single-GPU, 24-hour, at-most-20-round setting. Its plug-and-play framework allows tasks, training frameworks, and tools to be swapped, so the conclusions can be extended to other learning paradigms and data modalities; reusing trajectories as mid-training data offers a reproducible path for turning evaluation into training resources.
In four of six model-task combinations, prediction-enabled runs achieve lower best scores than standard-run means, with the largest deficits in knowledge injection; whether explicit forecasting imposes an additional burden and how prediction accuracy relates to optimization performance remain open questions. The trajectory-reuse experiment uses one 14B base model and a 5B-token budget, so whether gains scale with size and data mixture is unclear. The per-run improvement statistics include checkpoint selection and training variability and do not isolate the causal contribution of feedback itself; the out-of-distribution gaps in tool use and retrieval also invite further observation of their sources.
