AREX-2 trains a 27B agent on long-horizon reflective trajectories, reaching 81.8 on MLE-bench Lite and 84.0 on BrowseComp
Synopsis
AREX-2 defines agent self-improvement as turning more test-time rounds into a better solution, synthesizes long-horizon improvement trajectories from machine learning engineering and algorithmic programming tasks that retain failures and regressions, and after fine-tuning Qwen3.8-27B reaches 81.8 on MLE-bench Lite and 70.7 on Frontier-CS while lifting BrowseComp to 84.0, HLE to 52.6, GAIA to 92.2 and DeepSearchQA to 93.8 without adding new deep-research data, with continued gains as the round budget grows.
Interpretation
The paper separates self-improvement into two complementary capabilities: reflection sets how much a productive round is worth, and long-horizon execution sets how many rounds stay productive, with total improvement formalized as approximately their product. Prior agent systems implement retry-and-keep logic largely in the scaffold while the model produces one attempt at a time; this work moves the loop inside the model and defines self-improvement as test-time scaling within a single task. The claim rests on the paper's formalization and on its observation that current agent training data is built one attempt at a time, so it is a conceptual and framing argument.
A teacher model compiles GitHub repositories and online-judge problems into executable environments, and a teacher agent then works each task over hours and hundreds of tool calls, submitting and revising repeatedly to produce trajectories that contain failed runs, regressions and abandoned approaches. Trajectories are selected as a whole, requiring only a high final score and a compliant process rather than success in every round, so failures and regressions are deliberately kept as learning signal. Environment admission is checked by execution twice: the reference solution must run and score, and a simple baseline must score well below the reference, showing the environment leaves room for improvement.
AREX-2 reaches 81.8 on MLE-bench Lite, the highest in the comparison table, 8.1 points above the strongest baseline Naive-N0.5-Flash and 9.1 points above the strongest closed-weight baseline GPT-5.6 Sol, and 70.7 on Frontier-CS, 16.0 points above the strongest reported open-weight baseline. These results come from a 27B-parameter model while the compared systems include substantially larger ones. MLE-bench Lite reports Any Medal averaged over three seeds, Frontier-CS uses the official Agent Track setting with a 5-hour budget per task, and MLE-Lite follows the OpenMLE protocol with a 12-hour budget per task.
On four deep-research benchmarks AREX-2 reaches 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA and 93.8 on DeepSearchQA, surpassing both AREX (4B) and AREX (122B) trained with the previous recipe, with no new deep-research training data added. None of the newly constructed trajectories is a search task and the deep-research data is unchanged from the previous recipe, so the gains can be read as cross-domain transfer. The paper also reports a round-scaling curve on BrowseComp rising from 64.8 at 47 turns to 84.0 at 143 turns, with gain per turn about three times that of AREX at less than a quarter of its size.
Perspective
The work targets settings where a task can be attempted repeatedly and verifiable feedback is available, such as machine learning engineering and algorithmic programming; its data construction depends on domains like GitHub repositories and online judges that offer unambiguous feedback, room for sustained improvement and abundant source material. For deep research, where no external correctness signal exists, the paper shows the model can still improve by continuing to search and revise on its own judgment, provided the task allows a long round budget. The ablation shows operational knowledge (general and task-specific skills) plus more rounds lifts the base model from 28.8 to 68.2, and training adds a further 13.6 points under matched skills and rounds, so the approach is most valuable to deployers willing to spend inference budget and skill context.
The paper itself notes a remaining gap on the strongest BrowseComp and HLE results, so transfer does not reach frontier level on every benchmark. The round-scaling curves show diminishing returns over time, so the marginal value of a larger budget has to be judged per task. Trajectory selection depends on final score and process compliance, and the details of that compliance judgment are not expanded within this text. Skills also remain in context during MLE-Lite evaluation, and while the ablation separates the contributions of skills and training, how skills act at test time remains an open question. The authors propose widening training domains, lengthening horizons, and letting the model's own trajectories become its next training data.
