Skip to main content
Back to timeline
arXivSource publication:

Teaching a Builder to design execution environments: meta-skills lift macro-average scores by 8.95 points with frozen model weights

Synopsis

The work introduces meta-skills — principles specifying when support is needed, what resources to provide, and how the Target should use them — which a Builder learns from the Target's execution feedback on a development set while both models' weights stay fixed, then freezes into a skill bank to construct harnesses for unseen tasks; across Harness-Bench and NewtonBench, full-bank meta-skills raise macro-average performance by 8.95 percentage points over no-skill construction and 12.02 points over delivering the same bank directly to the Target.

AI-generated editorial illustration: Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Interpretation

The paper makes support design itself an object of learning: the Builder distills reusable meta-skills from the execution records and scores of the harnesses it built on development tasks, where each meta-skill has three fields — when (observable conditions calling for support), provide (the capability or resource to supply), and use (how the Target should employ it and which judgments remain its responsibility). Prior work learns task skills for the solver (e.g., Evo-Harness compiles execution experience into solver skills), searches harness implementations (e.g., Meta-Harness), or evolves strategies for engineering context files and code (e.g., Meta Context Engineering); this work focuses on reusable support principles at the Builder level and couples them with a deployment protocol that freezes the bank and rebuilds a harness per held-out task. Skill learning starts from an empty bank and runs two development-set passes, with at most one evidence-grounded keep/revise/add update per development task and a 192-token limit per skill; the development split is about 10% of each benchmark (Harness-Bench: 11 development and 95 test tasks; NewtonBench: 32 development and 292 test tasks), and the bank is frozen before testing.

Giving the same full meta-skill bank to the Builder to implement beats giving it directly to the Target: it wins in all six model–benchmark settings, by 12.02 points on average and up to 25.43 points in a single setting. Because the semantic knowledge content is held fixed and only the question of whether a Builder first turns it into executable support changes, the gap is attributed to operationalizing declarative principles as persistent state, executable tools, verification logic, or control decisions rather than merely exposing the Target to better advice. All six original settings favor the method, e.g., 67.81 versus 42.38 for Gemini on Harness-Bench; conditions share component interfaces, design guidance, construction permissions, and up to two interface-repair attempts, with Target execution budgets fixed.

With construction capability held fixed, experience itself adds value: the full-bank Builder outperforms the no-skill Builder in all six settings, by 8.95 points on average, reaching a 65.31% macro-average. The no-skill Builder keeps the same seven component families, construction instructions, and 16K output-token budget, so the gap reflects knowing what support to build and how to integrate it with Target behavior rather than scaffold-building ability alone. Gains vary by Target and benchmark: on NewtonBench all three Targets improve, averaging 10.96 points, versus 6.95 points on Harness-Bench; the authors read this as experience being most useful when failure modes recur consistently across tasks.

When the same model serves as both Builder and Target, meta-skills still help: across three same-model settings they average 18.71 points over no-skill construction and 14.14 points over direct delivery to the Target, with weights unchanged. This suggests a route to system-level self-improvement that needs neither a stronger external teacher nor weight updates: a model improves its own execution by learning to build better support for itself. Three settings use Gemini-3.6-Flash and Gemini-3.1-Pro, with construction, reflection, and execution performed by the same model; the authors note the evidence spans one main Builder, two benchmarks, and one adopted execution per task and condition.

Perspective

The results apply to a test-time AI4AI setting: Builder and Target weights are fixed, learning and construction happen outside the Target execution budget, and Target execution budgets are fixed (Harness-Bench: 30 turns, 30 tool calls, 96K tokens; NewtonBench: 12 turns, 10 tool calls, 192K tokens). The intended subjects are agent systems that must solve new instances within known workflow categories and physical mechanisms, with gains measured as macro-average score. For practitioners, this suggests distilling experience into a reusable bank of support principles that a Builder compiles per task into seven component families — instructions, memory, context organization, composed tools, execution control, verification and recovery, and workspace preparation. For researchers, it provides a reference point for comparing alternative support representations and construction policies.

Cross-Builder transfer is inconsistent in direction: on NewtonBench the Sol-to-Qwen bank gives Sol +13.36 points but −4.45 points when transferred to Gemini-Pro, while the same transfer gives +2.72 points on Harness-Bench, and all confidence intervals include zero. Skill refinement is not monotonic: on Harness-Bench Gemini improves with each pass while Qwen drops 4.06 points from its first-pass peak. In component ablations, only Gemini's controller removal has a confidence interval excluding zero, and all three memory-plus-context intervals contain zero. The authors also note the evidence spans one main Builder, two benchmarks, and one adopted execution per task and condition, leaving cross-dataset reuse, longer learning histories, and evaluation of gains relative to construction cost to future work.

Sources