Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

DUET co-evolves solver and grader agents in alternating rounds, improving both task execution and grading agreement across four agent benchmarks

DUET treats both a solver agent and a grader agent as optimization targets: each round it adaptively selects training tasks, the solver executes them, the grader scores the outcomes with evidence-grounded critiques, and a tool-using update module revises the optimizable component (system prompt or skill library) of either agent, alternating between solver and grader, with training updates using no reference answers or ground-truth labels; across SpreadsheetBench, GDPval, τ-Bench, and Finance Agent Benchmark, DUET improves both solver and grader performance and surpasses fixed-grader baselines SkillOpt and GEPA on solver metrics, even when those baselines use the benchmark reference evaluator.