Public articles linked to the same research event.
arXiv DUET treats both a solver agent and a grader agent as optimization targets: each round it adaptively selects training tasks, the solver executes them, the grader scores the outcomes with evidence-grounded critiques, and a tool-using update module revises the optimizable component (system prompt or skill library) of either agent, alternating between solver and grader, with training updates using no reference answers or ground-truth labels; across SpreadsheetBench, GDPval, τ-Bench, and Finance Agent Benchmark, DUET improves both solver and grader performance and surpasses fixed-grader baselines SkillOpt and GEPA on solver metrics, even when those baselines use the benchmark reference evaluator.
DUET treats both a solver agent and a grader agent as optimization targets: each round it adaptively selects training tasks, the solver executes them, the grader scores the outcomes with evidence-grounded critiques, and a tool-using update module revises the optimizable component (system prompt or skill library) of either agent, alternating between solver and grader, with training updates using no reference answers or ground-truth labels; across SpreadsheetBench, GDPval, τ-Bench, and Finance Agent Benchmark, DUET improves both solver and grader performance and surpasses fixed-grader baselines SkillOpt and GEPA on solver metrics, even when those baselines use the benchmark reference evaluator.
DUET treats both a solver agent and a grader agent as optimization targets: each round it adaptively selects training tasks, the solver executes them, the grader scores the outcomes with evidence-grounded critiques, and a tool-using update module revises the optimizable component (system prompt or skill library) of either agent, alternating between solver and grader, with training updates using no reference answers or ground-truth labels; across SpreadsheetBench, GDPval, τ-Bench, and Finance Agent Benchmark, DUET improves both solver and grader performance and surpasses fixed-grader baselines SkillOpt and GEPA on solver metrics, even when those baselines use the benchmark reference evaluator.
DUET treats both a solver agent and a grader agent as optimization targets: each round it adaptively selects training tasks, the solver executes them, the grader scores the outcomes with evidence-grounded critiques, and a tool-using update module revises the optimizable component (system prompt or skill library) of either agent, alternating between solver and grader, with training updates using no reference answers or ground-truth labels; across SpreadsheetBench, GDPval, τ-Bench, and Finance Agent Benchmark, DUET improves both solver and grader performance and surpasses fixed-grader baselines SkillOpt and GEPA on solver metrics, even when those baselines use the benchmark reference evaluator.