Skip to main content
Back to timeline
arXivSource publication:

Ant Group proposes Marathoner: synthesizing tasks from million-line-scale GitHub PRs lets a 9B open model work 10+ hours and make 1000+ tool calls

Synopsis

The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.

AI-generated editorial illustration: Marathoner: Ultra-Long-Horizon Autonomous Intelligence

Interpretation

Proposes a comprehensive post-training pipeline for ultra-long-horizon execution capability, comprising ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning. Previously, ultra-long-horizon execution was mainly demonstrated by proprietary models whose training methods and implementation details were rarely disclosed; this work instills the capability into an open base model through a reproducible pipeline. The paper provides a full pipeline description and implementation details, including data scale, training hyperparameters, frameworks, and hardware, plus comparisons on 5 benchmarks.

Synthesizes task-level data from major release PRs on GitHub and introduces Multi-Task Chaining to chain multiple tasks into harder Frontier tasks. Task difficulty is stratified by lines of new code (Easy 100-200, Medium 200-1000, Hard 1000+), with a 2:3:5 ratio; chaining 5 atomic tasks yields tasks requiring coordination across multiple large repositories. The paper reports collecting 10,000 GitHub repositories, mining 100,000 major release PRs, and chaining 15,000 tasks into 3,000 highly challenging tasks; ablations show chaining 5 tasks performs best.

Proposes Later Stage Bonus Reward, which explicitly rewards valuable operations during the later stages of execution in reinforcement learning. The reward targets the tendency to lose effective progress in the later stages of ultra-long-horizon execution, using an LLM to summarize the trajectory into phases and judge whether the latter half contains key operations, granting an additional 0.5 bonus if so. Ablations show a bonus of 0.5 covering the 50%-100% portion of execution performs best; smaller or larger bonuses degrade performance.

Marathoner-9B achieves consistent improvements over the base model across 5 ultra-long-horizon benchmarks and surpasses strong proprietary models on some. The base Qwen3.5-9B scores 10.2 on FrontierSWE and 0 on SWE-Marathon, while Marathoner-9B reaches 26.4 and 8.2 respectively; it surpasses Gemini-3.1-Pro on FrontierSWE and SWE-Marathon. Results come from a comparison table across 5 benchmarks, and the paper emphasizes the capability is developed through a unified training pipeline without benchmark-specific optimization.

Perspective

The result targets software engineering tasks requiring sustained long execution, applicable in settings with sandbox environments and verifiable unit tests; for researchers and engineering teams building open ultra-long-horizon agents, it offers a concrete path using GitHub PRs as task sources, unit tests as reward verification, and multi-harness sandboxes as execution environments. The execution statistics in the paper (FrontierSWE averaging 3.56 hours, 426.2 steps, 648.3 tool calls) and the case record of 11.8 hours, 872 steps, and 1,273 tool calls indicate the capability can be observed and measured within controlled sandboxes.

The paper reports results on 5 benchmarks, among which SWE-Marathon and FrontierSWE have relatively few tasks (20 and 17 respectively), so the stability of single-point differences needs more tasks to verify; Later Stage Bonus Reward relies on an LLM to phase trajectories and judge value, and the consistency of that judgment is not elaborated in the text; moreover, the safety and controllability of ultra-long-horizon execution in open environments is only mentioned in the introduction regarding risks of proprietary models, without an evaluation specific to this model.

Sources