Skip to main content
Back to timeline
arXivSource publication:

ActiveSaddler co-evolves training scenarios with the harness, raising Pass@1 by 4.4 and 7.5 points on GAIA2 and Terminal-Bench 2.0

Synopsis

The work formulates training-scenario selection in offline harness optimization as an automated curriculum-learning problem and introduces ActiveSaddler, which models the curriculum as a non-stationary bandit with failure-pattern arms, estimates each arm's potential learning progress, and adaptively balances revisiting known weaknesses against exploring unseen scenarios, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer with a scenario order fixed before optimization on GAIA2 and Terminal-Bench 2.0 respectively.

AI-generated editorial illustration: ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Interpretation

The paper observes that existing automated harness optimization methods mainly optimize how the harness is updated while the training scenarios that generate feedback are usually fixed in advance, either as the full training set or as mini-batches scheduled before optimization; as the harness evolves, which scenarios are most useful can change, so the curriculum itself should adapt. The authors formulate this missing dimension as an automated curriculum-learning problem for budgeted offline harness optimization and give a formal objective for a curriculum policy that selects the training-scenario set. The argument rests on the problem formulation (Equations 5 and 6) together with the observation that under fixed-order optimization many failures observed during training remain unresolved by the final optimized harness, making this a conceptual plus empirical motivation.

ActiveSaddler models the curriculum as a non-stationary bandit whose arms are instantiated online from observed failures: a Failure-Pattern Extractor abstracts failures into symptoms and normalizes them into reusable failure-pattern arms, an Arm Prioritizer uses an LLM to estimate a learning-progress score from severity, fixability, breadth, and side-effect risk, and an Exploration Controller decides whether to keep repairing a known arm or execute unseen scenarios. Unlike arms defined at category or single-scenario level, failure-pattern arms target a harness weakness rather than an individual failed execution, letting related failures across scenarios share one optimization target and letting the arm set expand during optimization. Method details appear in Section 4 and Appendices A, B, and K, including the Pattern Registry, the pattern CLI command set, Algorithms 1 and 2, and the three prompt templates; ablations show category and scenario arms reduce test Pass@1 by 3.0 and 3.6 points on GAIA2 and by 10.8 and 6.7 points on TB2.

On GAIA2 and Terminal-Bench 2.0, ActiveSaddler achieves the strongest test performance, improving over the same AutoSaddler optimizer with a randomly shuffled scenario order fixed before optimization by 4.4 and 7.5 percentage points respectively, and outperforming GEPA, Meta-Harness, AutoSaddler, and fixed curricula ordered by category or scenario accuracy. This is the first treatment of which training scenarios generate feedback as an adaptive dimension in harness optimization, compared directly against fixed curricula and existing optimizers under the same rollout budget. The main table reports a 300-task GAIA2 test split and a 40-task TB2 test split, with three test-time executions per optimized harness and mean Pass@1 with standard deviation; Appendix C shows ActiveSaddler beating all fixed-order baselines in two independent GAIA2 runs, and Appendix E shows applying ActiveSaddler to GEPA raises average Pass@1 from 54.2 to 57.2.

Ablations show the gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery; removing the Arm Prioritizer reduces GAIA2 and TB2 by 4.5 and 7.5 points, and replacing adaptive exploration with a fixed every-five-iterations schedule reduces them by 4.5 and 8.3 points. These ablations isolate the contribution of each curriculum component and compare against non-LLM alternatives, with EMA-based failure-persistence scoring and a count-based UCB-AIR controller falling 3.1 and 4.2 points below ActiveSaddler. Ablations run under the same optimizer and the same rollout budget; Appendix G reports failure hit rates of 52.9% for failure-pattern arms versus 36.4% for category arms and 17.6% for scenario arms on GAIA2, and 60.4% versus 30.5% and 36.6% on TB2, while Appendix J shows 2.0 and 6.9 times as many unresolved arms at Pull as at Draw decisions on GAIA2 and TB2.

Perspective

The work targets budgeted offline harness optimization under a standard train/dev/test protocol, and the curriculum layer is decoupled from the underlying optimizer, so it can be layered on offline optimizers such as AutoSaddler and GEPA; the beneficiaries are research and engineering teams that automatically tune prompts, tool interfaces, and runtime control logic for LLM agents. The authors explicitly describe this as a research exploration conducted entirely in benchmark environments with public and synthetic task data, with no real user data, and they do not evaluate production safety, security, privacy, or governance readiness, so the results characterize optimization effectiveness under controlled benchmark settings rather than clearance for production deployment.

Failure-pattern arms are induced and normalized by an LLM from execution failures, and how their granularity and stability shift across task distributions is illustrated through case studies and hit rates but remains a question worth tracking. Both the Arm Prioritizer and the Exploration Controller rely on LLM judgments, and Appendix D shows simple non-LLM alternatives perform worse, indicating these judgments carry real weight in the current setting; how they behave with stronger or weaker models is still to be examined. The cost analysis shows ActiveSaddler has higher optimizer-side overhead ($11.44 versus $7.87 per generated patch on GAIA2), with end-to-end advantage coming from rollout allocation efficiency, so whether this trade-off holds under different prices and budget structures needs further observation. In addition, this document is a full-text parse in which figures appear as text and tables, so some graphical detail, such as the full trajectory of probability mass across iterations, can only be understood indirectly through the prose and the numerical values in the appendices.

Sources