Skip to main content
Back to timeline
arXivSource publication:

Regularizing the Self-Improvement of Agent Harnesses: How RRSI Makes Recursive Improvement Transfer

Synopsis

The work introduces RRSI, which keeps an agent harness (prompts, control flow, tooling, memory and context management) fully editable while regularizing both the proposal and the selection of edits, gaining up to 14.1 points on the split it evolves against and up to 4.7 points on five out-of-distribution benchmarks across eight benchmarks in coding, agentic workspace and engineering design, while running on 30% fewer policy tokens than unregularized evolution.

AI-generated editorial illustration: RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Interpretation

It identifies overfitting as a key challenge in harness-based recursive self-improvement, tracing it to coupled behaviors: benchmark-specific fitting, chasing evaluation noise, and complexity accumulation. Prior harness-evolution methods are driven by the score on the suite they optimize against; this work makes generalization explicit and defines it as an evolved harness transferring to unseen benchmarks with different task descriptions, tool interfaces or verifiers. The paper frames the problem and analyzes mechanisms, drawing on recent studies that report gaps between evolution and held-out performance and on delta attribution that separates reusable mechanisms from task fitting.

It proposes RRSI, which constrains the search on the proposal side with an annealed edit budget, evidence-aware credit assignment and structured exploration, and on the selection side with leakage screening, a noise-adjusted floor and complexity-aware acceptance. Rather than restricting which harness components may change, RRSI keeps the edit space open and regularizes the search trajectory; the authors describe the edit budget as analogous to a cardinality constraint, structural pruning as Lasso-style sparsification, and complexity-aware acceptance as Ridge-style shrinkage, noting these are role analogies rather than norm-penalized objectives. The method is specified with a round-level formulation, atomic edit representation, edit history and component-level summaries, noise floor and acceptance conditions, plus a hyperparameter table per instance.

Across eight benchmarks in three domains, RRSI improves the evolve split and every held-out split, with no held-out split regressing. Evolve-set gains are 6.0 points on Terminal-Bench 2.1, 4.9 on EngDesign and 1.1 on Harvey LAB; transfer gains are 1.8 points on SWE-bench Verified, 2.3 on the in-distribution Harvey LAB held-out split, 3.5 to 4.7 points (7.2% to 13.1%) on three out-of-distribution agentic benchmarks, and 4.3 Medal points (24.3% relative) on Frontier-Eng. Every number is measured against the unevolved base harness in the same window with the same tool environment, judge and number of trials; Harvey LAB's 160 tasks are partitioned once into a 120-task evolve set and a 40-task held-out set and fixed.

The regularization buys both transfer and cost: RRSI produces the lightest harness of any evolved harness compared, and ablations show that removing either group of regularizers raises the evolve-set score while lowering transfer. Removing acceptance constraints raises the evolve score from 90.5 to 91.5 while the out-of-distribution average falls from 43.6 to 41.0 and token cost rises by half; removing proposal constraints costs only 0.2 points on the evolve split but 1.7 out of distribution; removing both lifts the evolve score to 92.8, the highest of any arm, leaves the out-of-distribution average at 40.3, near the unevolved harness, at 3.80 million tokens per trial versus 2.42 for RRSI. The ablation runs on the agentic workspace instance with shared base harness, policy, evolve split, round count and candidate budget; cost is measured in policy tokens per trial and trajectory steps, with RRSI at 26.3 steps versus 27.3 to 34.6 for prior methods.

Perspective

The results apply to settings with a frozen backbone where the harness is the optimization variable, across coding, agentic workspace and engineering design tasks, and transfer is shown across two policy families and a smaller backbone that never took part in the search; for teams automating harness engineering while controlling inference cost, it offers a reusable constrained-search recipe.

The authors note that the study does not address settings where model weights are updated during evolution, that RRSI still relies on a finite evolve set and several regularization hyperparameters, and that effectiveness may depend on feedback-signal quality and the chosen search budget; broader validation across substantially different agent architectures, tool ecosystems and longer-running self-improvement processes remains open. In addition, this reading is of the full text, but some figure and appendix values appear in textual form, so exact experimental details are best checked against the original and the round-by-round records on the project website.

Sources