Public articles linked to the same research event.
arXiv The work presents a meta-agent pipeline in which a frontier model such as Claude Opus generates terminal tasks and verifiers, diagnosing three failure classes—benchmark invalidity, harness brittleness, and reward misalignment—and reports that prompt redesign and context extension raise baseline solvability 5.6 times, that a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus-generated tasks, and that adding hard tasks reduces mean pass@2 to 20.6% without changing the training configuration, indicating that the solvability band is model-specific.
The work presents a meta-agent pipeline in which a frontier model such as Claude Opus generates terminal tasks and verifiers, diagnosing three failure classes—benchmark invalidity, harness brittleness, and reward misalignment—and reports that prompt redesign and context extension raise baseline solvability 5.6 times, that a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus-generated tasks, and that adding hard tasks reduces mean pass@2 to 20.6% without changing the training configuration, indicating that the solvability band is model-specific.
The work presents a meta-agent pipeline in which a frontier model such as Claude Opus generates terminal tasks and verifiers, diagnosing three failure classes—benchmark invalidity, harness brittleness, and reward misalignment—and reports that prompt redesign and context extension raise baseline solvability 5.6 times, that a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus-generated tasks, and that adding hard tasks reduces mean pass@2 to 20.6% without changing the training configuration, indicating that the solvability band is model-specific.
The work presents a meta-agent pipeline in which a frontier model such as Claude Opus generates terminal tasks and verifiers, diagnosing three failure classes—benchmark invalidity, harness brittleness, and reward misalignment—and reports that prompt redesign and context extension raise baseline solvability 5.6 times, that a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus-generated tasks, and that adding hard tasks reduces mean pass@2 to 20.6% without changing the training configuration, indicating that the solvability band is model-specific.