DUET co-evolves solver and grader agents in alternating rounds, improving both task execution and grading agreement across four agent benchmarks
Related research and updatesSynopsis
DUET treats both a solver agent and a grader agent as optimization targets: each round it adaptively selects training tasks, the solver executes them, the grader scores the outcomes with evidence-grounded critiques, and a tool-using update module revises the optimizable component (system prompt or skill library) of either agent, alternating between solver and grader, with training updates using no reference answers or ground-truth labels; across SpreadsheetBench, GDPval, τ-Bench, and Finance Agent Benchmark, DUET improves both solver and grader performance and surpasses fixed-grader baselines SkillOpt and GEPA on solver metrics, even when those baselines use the benchmark reference evaluator.
Interpretation
DUET turns the grader from a fixed source of feedback into a first-class optimization target alongside the solver, so task execution and evaluation evolve together through alternating updates. Prior agent optimization methods such as GEPA and SkillOpt keep the grader fixed while optimizing the solver, whereas DUET updates the grader inside the optimization loop. Across four benchmarks DUET achieves stronger solver performance than fixed-grader baselines; on Finance Agent Benchmark DUET reaches 96.5% rubric score versus 92.8% for the unoptimized solver, 94.0% for SkillOpt, and 94.4% for GEPA, and also exceeds SkillOpt (95.7%) and GEPA (95.9%) when they optimize against the benchmark reference evaluator.
Co-evolution also improves the grader's own evaluation capability. The grader develops more reliable scoring and critique from a minimal initialization, whereas the baselines do not produce a trained grader. On Finance Agent Benchmark grader pairwise agreement improves from 67.9% for the initial grader to 75.0%; on GDPval it improves from 68.4% to 78.5%, a 10.1 percentage point gain; on τ-Bench and SpreadsheetBench, whose reference evaluators give only binary outcomes, agreement changes only slightly (88.7% to 89.3% and 71.9% to 72.8%).
Three design elements—alternating updates, adaptive task selection, and bounded component updates—each contribute to DUET's performance. Ablations tie these design choices to overall performance rather than reporting only final results. On Finance Agent Benchmark adaptive task selection improves the solver from 95.9% to 96.5% and grader agreement from 62.7% to 75.0%; alternating updates improve the solver from 95.4% to 96.5% and grader agreement from 70.1% to 75.0%; optimizing only the grader with a fixed solver raises agreement only from 67.9% to 68.8%.
DUET extends beyond system prompts and remains robust across underlying models and prompt initializations. DUET can jointly optimize a system prompt and a skills library, and improves both agents under empty, informed, and adversarial initializations. On τ-Bench joint prompt-and-skill optimization reaches 91.9% success rate versus 89.4% for prompt-only DUET, 83.3% for SkillOpt, and 84.4% for GEPA; on Finance Agent Benchmark the adversarial solver improves from 91.1% to 95.8% and the informed solver from 93.3% to 95.0%, with DUET removing the harmful adversarial instructions in the first update.
Perspective
The framework targets open-ended agent tasks; training updates use no reference answers or ground-truth labels, and by default a labeled validation set is used only for checkpoint selection and early stopping, with a fixed-budget variant without validation-based selection also evaluated. It applies to optimizable components expressible as system prompts, skill libraries, or a combination, while underlying models, tools, and execution environments stay fixed. Results cover four benchmark families—spreadsheet manipulation, multi-turn customer-service tool use, and retrieval-augmented question answering—and are evaluated across different underlying models and empty, informed, and adversarial prompt initializations. For teams that want to bootstrap evaluation on a new task collection without ready-made grading rubrics or reference answers, this setting indicates that evaluation practices can develop alongside solver improvement.
Grader evaluation relies on benchmark reference evaluators rather than human judgments, which the authors list as a scope limitation; on τ-Bench and SpreadsheetBench the reference evaluators provide only binary outcomes, so preference pairs mainly capture coarse pass–fail differences and subtler grader judgment improvements may not be captured by these benchmarks. In the hyperparameter ablations, differences for the exploration coefficient are modest relative to observed variance. In addition, some tables and equations in the loaded text are not fully rendered, so certain ablation and cost comparisons rest on the prose rather than exact numbers; readers needing precise figures should consult the original tables.
