AutoDataBench makes agents deliver training tasks one at a time: five frontier agents all score below 20 out of 100 at 45 minutes, with difficulty calibration rather than mode coverage as the binding constraint
Synopsis
The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.
Interpretation
The paper makes the acceptance of a single synthesised task the object of evaluation, rather than using post-training model performance as a proxy. Earlier evaluations either ask for a trained checkpoint (as PostTrainBench does) or measure environment generation through post-training gain (as RSIBench-Data does while still scoring a checkpoint), whereas the data industry delivers sample by sample and accepts each sample against criteria; AutoDataBench judges before any training and at the granularity of one task. The paper compares AutoBencher, BenchAgents, InnovatorBench, PostTrainBench, RSIBench-Data and AutoDataBench property by property on artifact delivered, weakness-targeted, per-artifact verdict, difficulty in a band, no training run, and domains covered; the unit of evaluation is an episode in which the author agent delivers exactly one task directory.
At the default 45-minute budget, five frontier author agents all score below 20 out of 100, and the loss comes mainly from difficulty calibration rather than mode coverage. The paper reports kimi-k3 at 0.184, gpt-5.6-sol at 0.177, qwen3.8-max at 0.139, glm-5.3 at 0.130 and deepseek-v4-pro at 0.097; the share of deliveries landing in the pass-rate band ranges from 14.6% to 25.0%, while coverage of the hidden rubric among gate-passing deliveries is high, from 0.700 to 0.967. Each of the 24 original tasks, eight from each of three suites, is authored twice at the default budget, the target model is deepseek-v4-pro throughout, and figures are means over those two episodes; the paper states this split between high coverage and poor difficulty calibration is the opposite of the failure it expected.
Raising the strongest agent's time budget from 45 to 180 minutes lifts its score from 0.184 to 0.552, while the cost of one usable task stays almost unchanged. The paper reports all three suites improving in the same direction, with Terminal-Bench, the weakest at the shorter budget, gaining the most; the share of in-band deliveries that clear the gate rises from 0.900 to 0.960, so the agent is not reaching the band by restating the original; unit cost is $11.88 against $12.62 and 244 minutes against 230. The comparison holds the target model, the rubrics, the judge and the scoring fixed, but the paper states explicitly that the 180-minute arm also ran at the agent's maximum reasoning-effort setting, so two variables move together and the result is read as a claim about direction rather than as a measurement of the factor.
Author-agent rankings are not stable across domains, and cost ordering disagrees with score ordering. kimi-k3 is first on AutomationBench and fourth on Terminal-Bench; qwen3.8-max scores zero on AutomationBench and is first on Terminal-Bench; glm-5.3 scores zero on Terminal-Bench-Science. One usable task costs between $4.81 and $27.69, and kimi-k3 and gpt-5.6-sol score within 0.01 of each other while one costs several times the other. The paper gives a per-suite score figure and a per-agent table of minutes and dollars, and notes that each suite contributes only eight tasks, so per-suite numbers should be read as estimates over that fixed subset.
Perspective
The benchmark addresses data production for executable agent tasks: a task must be delivered in the suite's native directory layout with an instruction, an executable environment and a verifier, and the author agent may read the suite's conventions, the original task and the target model's record but may call only the target model. It applies to pipelines that accept data sample by sample and judge usability before training, covering terminal work, software engineering, scientific computing and cross-application business workflows. The paper deliberately narrows the job to writing one new task of similar structure for the same suite, so the results speak to that simplest form of data production rather than to inventing a domain, a format or quality standards from scratch. For teams using agents to author training data, the practical implication is to choose the author agent per domain rather than once, and to treat the time budget as an adjustable resource.
The assumption underneath the quality term remains untested: the paper states explicitly that whether a task exercising a measured mode yields training data that repairs it is still open. In the time-budget comparison the 180-minute arm also ran at maximum reasoning effort, so two variables move together and the paper reads that result as a claim about direction. Per-suite numbers rest on a fixed subset of eight tasks per suite and should be treated as estimates over that subset. The loaded text is the full paper with appendices, but figure values appear through the prose and tables; checking figure-level detail still requires the original.
