Skip to main content
Back to timeline
arXivSource publication:

MILO co-evolves agent harnesses and their search strategy, reaching 86.1% on Terminal-Bench 2.1 above the official leaderboard top and tightening three mathematical bounds

Synopsis

MILO (Meta-evolutionary Island Orchestration) co-evolves an agent harness together with the strategy that discovers it: a hierarchical island-based lineage memory treats rejected mutations as negative evidence, per-island mutator agents rewrite complete harnesses, and an orchestrator adapts the search through lineage grafting, speciation, mutator reassignment, and curriculum revision; across Terminal-Bench 2.1, PaperBench, and DeepSWE it improves resolution over its initial harness by 12.0%, 28.3%, and 10.3% with Opus 4.8 (versus best prior-search gains of 4.5%, 18.3%, and 0%), reaches 86.1±2.0% on Terminal-Bench 2.1 (official leaderboard top 83.8±2.

AI-generated editorial illustration: MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Interpretation

MILO casts harness discovery as strategy-level evolutionary search whose search strategy rewrites itself as progress stalls: an island-based hierarchical lineage memory keeps every seed-to-candidate path including rejected candidates, mutator agents combine global search history with parent-specific weakness feedback to rewrite whole harnesses, and an orchestrator applies Graft, Speciate, Reassign, and Curriculum interventions at plateaus. Prior methods either optimize only prompts or skills (GEPA, A-Evolve) or rewrite the whole harness while inheriting a fixed, exploitation-biased LLM mutator (Self-Harness, Meta-Harness), and most retain only survivors (a single candidate, a Pareto front, or the fittest per MAP-Elites cell), so failed directions are prone to being retried. The paper formalizes the three components (hierarchical lineage memory, evidence-driven mutators, meta-evolutionary orchestrator) and reports an additive ablation ladder (flat-memory LLM mutator, mutator agent, lineage memory, multiple islands, orchestrator), describing what each added mechanism changes; the per-rung ablation numbers are not fully rendered in the loaded text.

Across three long-horizon benchmarks, MILO-discovered harnesses outperform eight state-of-the-art harnesses and six search methods while raising accuracy and lowering inference cost. With Opus 4.8 the paper reports resolution-rate gains over the initial harness of +12.0%, +28.3%, and +10.3%, against best prior-search gains of +4.5%, +18.3%, and 0%; on Terminal-Bench 2.1 it reaches 86.1±2.0%, above the official leaderboard's top entry of 83.8±2.3%, while consuming 26% fewer tokens than its initial harness. All search methods share the same 72-hour wall-clock limit, the same three expert-designed seeds (B10–B12), and one shared scorer; error bars are task-clustered 95% confidence intervals, and PaperBench uses the official rubric-weighted replication score from a GPT-5.5 judge. Most cells of the results tables are empty in the loaded text, so per-cell numbers rest on the abstract and prose.

MILO's harnesses transfer across benchmarks and backbones, generalizing to the harder Frontier-Bench without further search. The paper stresses that harness gains need not transfer across backbones: on TB2.1 all eight state-of-the-art harnesses beat the minimal harness with Opus 4.8, but four underperform it with gpt-oss-120b; MILO reports gains on both backbones. Transfer evidence comes from evaluation on Frontier-Bench without additional search and from comparisons across the Opus 4.8 and gpt-oss-120b backbones; the paper also gives a concrete counterexample where Mini-SWE-Agent uses nearly Goose's tokens for comparable RR@5 on TB2.1 (73.9% vs. 74.5%).

On instance-level scientific discovery, MILO-evolved harnesses tighten the best-known upper bounds for three open mathematical problems. The paper reports Erdős minimum-overlap improving from 0.3808586 to 0.3808568, the first autocorrelation inequality from 1.50274365 to 1.50274360, and the third from 1.45081 to 1.44889, each below the prior best including AlphaEvolve, TTT-Discover, EvoX, and the live arena leader. In this experiment the agent edits optimizer programs, each task pairing a starting construction, a parent optimizer, a seed, and a 60-second run budget; harness fitness is the mean gain over a 30-task bank rescored by trusted subprocesses, and the paper notes that only AlphaEvolve's final program for the first autocorrelation inequality is public, so the other comparisons are at the construction level.

Perspective

The work targets agent developers and researchers who need long-horizon autonomous execution: given a runnable build_agent entry point with editable source, MILO discovers a stronger harness within a 72-hour wall-clock budget, applied to command-line tasks, paper replication, multi-file software engineering, and open mathematical problems that require editing optimizer programs. The paper instantiates its harness space on the DeepAgents framework but notes the method needs only a runnable build entry point and editable source, so it could in principle move to interaction schemes such as OpenHands or Mini-SWE-Agent; it also proposes extending to other solver families such as preference-optimization algorithms and design heuristics, and to instance-level discovery where automated harness discovery is applied to the agent that writes solution programs.

In the loaded text many numeric cells of the results tables are empty, so beyond the figures stated in the abstract and prose, per-benchmark scores, per-rung ablation gains, and the best harnesses of each baseline cannot be verified from the text; the paper also reports the best candidate from a single 72-hour search budget, so stability across backbones and benchmarks, and the reproducibility of the mathematical bound improvements on a larger problem set, remain directions a careful reader would keep watching.

Sources