Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Meta-agent-generated terminal tasks let a 9B model reach 81.3% mean pass@2 within 20 steps, dropping to 20.6% after hard tasks are added

The work presents a meta-agent pipeline in which a frontier model such as Claude Opus generates terminal tasks and verifiers, diagnosing three failure classes—benchmark invalidity, harness brittleness, and reward misalignment—and reports that prompt redesign and context extension raise baseline solvability 5.6 times, that a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus-generated tasks, and that adding hard tasks reduces mean pass@2 to 20.6% without changing the training configuration, indicating that the solvability band is model-specific.