Harness-Zero distills an evolved agent harness into Qwen3.5-9B weights, lifting macro-average success from 23.3% to 44.3% and beating the 41.7% of the base model with the harness still attached
Synopsis
The work introduces Harness-Zero: an agent-as-harness setup in which a harnessing agent reviews and minimally rewrites the student agent's proposals at its response boundary, translating guidance from an evolved harness (tools, middleware, skills, memory) into executable supervision inside the target harness's action space, followed by SFT on the reviewed trajectories; across SpreadsheetBench Verified, AppWorld, and USPTO retrosynthesis, training-free agent-as-harness averages 81.1% over six benchmark-model settings versus 78.1% for code-as-harness, while the distilled Qwen3.5-9B under a minimal target harness alone raises the macro average from 23.3% to 44.3%, surpassing the 41.7% of the base model with the evolved harness still attached, and recovers 82.
Interpretation
The paper formulates agent harness distillation as a problem: harness optimization improves the agent's external scaffolding rather than the model, so gains stay tied to that harness at deployment, while the best harness varies across domains, instances, and base models; a shared harness forgoes specialized gains and maintaining many harnesses requires routing and recurring costs in context, model calls, tool calls, and orchestration. Prior lines such as automated harness optimization (e.g., Meta-Harness) and model-harness co-evolution let harness and weights reinforce each other, but gains remain coupled to the harness; Harness-Zero explicitly targets moving harness-induced behavior into model parameters so that only a fixed target harness remains at deployment. The problem statement comes from the introduction and related work, which cites a controlled coding-agent study: changing the evaluation harness affected performance more than the training method, and training with feedback collected across harnesses did not improve transfer to a held-out minimal ReAct harness.
The core method is agent-as-harness: a harnessing agent intercepts each proposed response at the student's response boundary, uses a private reference harness (adapted from the evolved harness) to decide whether to intervene, passes sound proposals unchanged, and otherwise makes the smallest coherent replacement valid in the student's action space; the accepted response is executed through the target harness, its observation enters the student-visible trajectory, rejected proposals and private review discussion stay outside it, and SFT is applied to the accepted responses. Unlike code-as-harness, which wraps the student in code, agent-as-harness translates harness guidance into supervision expressed as student-native actions, bridging differences in action space and available information between source and target harnesses; adaptation turns tools into action recipes, student-side middleware into review middleware, and skills and memory into diagnostic criteria and intervention guidance. The paper gives the three-stage pipeline, the review protocol (one pass/replace submission per update, replacements must be complete responses, per-trial replacement budgets of 5, 3, and 1 for SpreadsheetBench and AppWorld and 5 for USPTO), and trajectory filtering rules (dropping trajectories containing private artifacts such as /components/ paths or the review tool name, and masking reviewer-perspective reasoning spans).
In training-free comparisons, agent-as-harness with the adapted reference harness averages 81.1% across six benchmark-model settings, above 78.1% for meta-harness and 68.6% for mini-SWE-agent; with an empty reference harness it averages only 69.2%, so review alone explains little of the gain. The paper reports a +27.6% average relative improvement over code-as-harness and argues that a code harness encodes assumptions about how a model should act that go stale as capabilities change, whereas agent-as-harness moves that adaptation into inference. Table 1 covers DeepSeek-V4-Pro and GPT-5.6 Sol on SpreadsheetBench, AppWorld, and USPTO; in both agent-as-harness conditions the same model plays student and harnessing agent, so gains cannot come from a stronger supervising model. Appendix F adds that the relative advantage correlates with model capability: +1.0% on GPT-5.6 Sol, +4.0% on DeepSeek-V4-Pro, -1.0% on DeepSeek-V4-Flash, and -12.0% on Qwen3.6-35B-A3B.
In the distillation experiments, Harness-Zero raises Qwen3.5-9B's macro average from 23.3% to 44.3% (+21.0 points, +90.1% relative) under the minimal target harness alone, exceeding the 41.7% of the base model with the evolved harness still attached; behavioral analysis shows the distilled model recovers 82.3% of 28 harness-exclusive patterns on average. Ablations show the supervision source matters more than stronger demonstrations: direct fine-tuning on GPT-5.6 Sol trajectories leaves the student at 12%, teacher trajectories under the evolved harness fall to 3%, student trajectories under the evolved harness return to 12%, review with an empty reference harness reaches 11%, and review with oracle answers reaches 15%, while Harness-Zero reaches 30%; collection success does not predict distillation value (the oracle condition succeeds on 98.6% of collection tasks yet yields 15%). Tables 2 and 3 both start from Qwen3.5-9B, use the same 500-task collection split and training recipe, and evaluate the 2-epoch checkpoint under the target harness; Table 4 recovery rates are scored only on support tasks where the base model exhibits the pattern under the evolved harness and never under the target harness, so the base model scores 0% by construction.
Perspective
The result targets a student model deployed under a fixed minimal target harness (mini-SWE-agent style with a single Bash execute tool) and applies to settings where domain-harness procedural behavior should be internalized into weights; the paper states this does not remove the need for a harness but narrows what that harness must provide, and expects a minimal one to suffice as harness distillation improves. The method requires a one-time harness evolution (here three rounds of skill-guided evolution run by Kimi K3 under Kimi Code) plus harness adaptation, and a sufficiently capable harnessing model to collect trajectories; collection costs at least two model calls per step, and on USPTO agent-as-harness takes 237.2 s per trial on average versus about 100.1 s for mini-SWE-agent, an overhead absent after distillation. What a reader can reuse directly is the review-minimal-replacement-SFT pipeline and the component adaptation mapping (tools to action recipes, middleware to review middleware, skills and memory to diagnostic criteria and failure patterns), plus the practice of translating harness rules into trajectory detectors to measure internalization.
Open questions the paper itself raises include: the method depends on a capable harnessing model, and review can be harmful with weaker models (in Appendix F, Qwen3.6-35B-A3B is -12.0% relative to meta-harness while replacing 66.0% of reviewed steps to net harm); SFT may not fully internalize deep domain knowledge encoded by the evolved harness, as on USPTO the distilled model reaches 30.0% versus 38.0% with the harness attached; and some harness mechanisms such as context management are not directly expressible as student responses. Future directions include the counterfactual prediction problem facing a harnessing agent, which could motivate training stronger agent world models; the current design reviews every proposal and only passes or replaces whole responses, so proxy signals could trigger review selectively and finer-grained mechanisms such as token insertion and latent-space steering could be supported; and each replacement pairs a rejected and a preferred response at the same state, so preference learning could use comparison signals that response-level SFT discards. A careful reader should also note that the distillation experiments center on Qwen3.5-9B and report single-run pass@1, so stability across base models and domains still needs more evidence; this is a full-text read, and figure and table values are taken as reported in the main text and appendix tables.
