AED turns 50,228 failed agent rollouts into error–diagnosis pairs: first-proposal corrections lift verifier pass rates from 18.4% to 51.1%, and full-diagnosis fine-tuning raises Qwen3-8B's exact-step agreement from 47.2% to 63.6%
Synopsis
The work introduces the Agent Error Dataset (AED), comprising 50,228 error–diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families and 23 policy models, and a five-stage Agentic Error-to-Training (AET) pipeline that links natural failures to diagnoses, proposed corrections and optional replay evidence; across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1% (a 32.7 percentage-point gain), full-diagnosis fine-tuning on a separately frozen diagnosis release raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6% (averaged over three seeds on a 943-case holdout), and in a single-seed actor-training comparison action-only repair training scores 6.
Interpretation
AED turns the observation that a failed rollout carries more than its final reward into a reusable, scaled resource: 50,228 error–diagnosis pairs spanning 33 environments, 19 harness families and 23 policy models, retaining source traces and execution metadata so that cross-setting failure analysis and re-diagnosis can proceed without repeating the original rollout. Prior failure-attribution resources were smaller and differently scoped, such as Who&When with 184 natural multi-agent failures, MAST with 1,642 traces and AgenTracer with 2,000 records, each targeting different objects (failure modes, responsible agents, attribution models); AED places natural failures, explanations, proposed corrections and execution evidence in one collection and supports both diagnosis and action training views. Collection counts come from the same source-linked pair index as Figure 3 and require at least ten distinct source tasks per environment, harness family or policy model; the paper states that the collection counts error–diagnosis pairs, not independent executions or training-ready examples, and that additional diagnoses of one execution add pairs rather than independent failure observations.
The AET pipeline links diagnoses and proposed corrections to controlled replay: across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a 32.7 percentage-point gain (task-clustered interval reported in the appendix). Earlier work largely stopped at failure attribution or injected-fault verification; AED compares a corrected action against an original-action retry from the same checkpoint under matched policy, harness, budget and verifier settings, and retains both successful retries and failed corrections so that passing branches are distinguished from evidence for an action preference. Paired replay uses temperature 0.2 for the first sample and 0.8 for subsequent samples, shared by both arms; the paper notes that original-action retries also succeed on some pairs, that the net gain comes from the difference between discordant cells, and that a single paired outcome does not identify an individual causal effect or prove a repair unnecessary.
Full-diagnosis SFT on a separately frozen diagnosis release raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6% (averaged over three seeds on a 943-case holdout), with mean agreement improving at each of four increasing training-set sizes; the strongest prompted reference in this comparison scores 54.7%. This provides an operational route to training a failure-diagnosis model from failed experience and reports a scaling ladder; the paper stresses that the comparison measures agreement with the recorded annotations, including their conventions and defects, rather than a general ranking of diagnostic ability, and that the student learns from consensus-generated labels while the references receive no training on those labels. Three seeds, a 943-case holdout and four nested training sizes; the paper also reports that public transfer is protocol-dependent: on Who&When the interval includes zero after excluding flagged task overlap, and mean TrajErrBench accuracy remains below base, so internal localization is the supported result while public transfer remains a scope limit.
In a single-seed comparison of actor-training recipes, action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training, while repair-containing recipes score lower on TextQuest; on real ALFWorld and ScienceWorld the preventive-repair arm loses to the untrained policy. This extends learning from failure beyond diagnosis into acting policies and separates preventive, post-error action-only and post-error reflective supervision; the paper notes that these recipes differ in task pools, exposure and optimizer updates, so their contrasts measure the combined recipe rather than repair supervision alone. Single seed across six development environments; after a post-hoc Holm adjustment the WebShop-lite action-only and reflective gains and all three TextQuest losses survive. The real-environment comparison covers only the preventive-repair arm, and ScienceWorld omits six run errors from its repair denominator.
Perspective
The resource targets text-based agent systems and is meant for cross-setting failure analysis, diagnosis training and actor recovery training; replay evidence holds only on the environment subset that supports replay and requires the same checkpoint, harness, budget and verifier. For researchers wanting to reuse failed experience, AED provides source traces, execution metadata and optional replay branches so that re-diagnosis does not require repeating the original rollout; for teams wanting to train diagnosis models or repair policies, it offers diagnosis SFT, recovery SFT and action-preference views with their own admission rules. The paper states that collection size counts error–diagnosis pairs and that learning results concern smaller frozen populations, so the resource is a starting point for failure analysis and error-aware post-training rather than a measure of full-collection utility.
Public-benchmark transfer is protocol-dependent: on Who&When the interval includes zero after excluding flagged task overlap, mean TrajErrBench accuracy remains below base, and AgentErrorBench point estimates improve without a clear paired advantage. Actor comparisons are single-seed, and the recipes differ in task coverage, exposure and optimizer updates, so their contrasts measure the combined recipe rather than repair supervision alone; on real ALFWorld and ScienceWorld only the preventive-repair arm has a comparison, and it declines. Diagnosis fine-tuning measures agreement with internal teacher labels generated by heterogeneous consensus, including their conventions and defects, so it is not a capability ranking on independent labels. Replay evidence depends on state restoration and harness support, and contrasts under a substituted harness or partly restored state do not transfer directly to the original system. In addition, this load is the full text, and the specific numeric tables in figures and appendices are not rendered cell by cell, so exact intervals and per-environment results still require consulting the original tables.
