Skip to main content
Back to timeline
arXivSource publication:

Mid-Harness verifies candidate actions between model and harness, lifting TMAX-9B Pass@1 on TerminalBench-Lite from 50.00% to 68.03%

Synopsis

The work introduces Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before forwarding one for execution while keeping the generator and harness unchanged, and studies action-level test-time compute scaling: on TerminalBench-Lite a GPT-5.6 Sol verifier raises TMAX-9B Pass@1 from 50.00% to 68.03% with 8 sampled actions, wider sampling yields little benefit under weak verification, pairwise verification performs best among the evaluated self-verification mechanisms, distilling the stronger verifier's responses into TMAX-9B further improves Pass@1, and combining action scaling with trajectory scaling (Best-of-N and Sequential Refine) reaches higher success at lower estimated token cost.

Interpretation

Verification quality governs the benefit of action sampling: with the TMAX-9B generator fixed on TerminalBench-Lite, zero-shot listwise verification changes Pass@1 only from 49.32% to 51.02% and Pass@3 from 66.33% to 67.35% when candidate width doubles from 4 to 8, whereas a GPT-5.6 Sol verifier under the same listwise mechanism reaches 64.63% at N=4 and 68.03% at N=8. Prior work allocated test-time compute to sampling, verification, and refinement but left an incomplete understanding of when and why this scaling improves trajectory success; this work jointly compares candidate width, verification mechanism, and verifier capability at a fixed model-harness boundary. The main evaluation uses 98 TerminalBench-Lite tasks with three runs per task and reports Pass@1 and Pass@3; the authors note the evaluation lacks gold action labels, so candidate coverage is probed indirectly through trajectory success under a strong verifier.

In the self-verification setting, pairwise verification performs best among the evaluated mechanisms, and LoRA distillation on 117k pairwise responses from GPT-5.6 Sol raises Pass@1 from 54.76% to 57.14% and Pass@3 from 71.43% to 75.51% at N=8, and Pass@1 from 54.42% to 55.44% at N=4, while leaving the action generator unchanged. The work treats verification mechanism (listwise, pointwise, pairwise) and verifier distillation as separable variables and quantifies the offline agreement shift: score MAE falls from 2.59 to 1.05, pairwise agreement rises from 59.01% to 74.58%, and verification agreement rises from 38.52% to 57.79%. Distillation uses 244 training tasks and 117,631 verifier inputs, with analysis on 21 disjoint held-out tasks and 10,543 inputs; the offline diagnostics measure agreement with the teacher verifier, which the authors state is not action correctness or trajectory success.

Action scaling and trajectory scaling are complementary: on TMAX-9B, using Mid-Harness to generate the source trajectories for Best-of-N raises Pass@1 from 55.10% to 61.22% (zero-shot) and 66.33% (distilled), an 11.23 percentage-point gain with the same three environment executions compared with Best-of-N alone, while distilled Mid-Harness with one round of Sequential Refine raises Pass@1 from 55.10% to 60.20% and Pass@3 from 71.43% to 75.51%. Prior trajectory-level methods (parallel verification and sequential refinement) require fresh environment runs; this work shows that inserting verification before action execution improves those methods without increasing their number of environment runs and reports the cost-success trade-off. Composition experiments run on TerminalBench-Lite at TMAX-4B, 9B, and 27B and report reference-priced token cost; the authors give task-level bootstrap 95% confidence intervals, with distilled 9B (+7.14, [+1.36, +12.93]) and distilled 27B (+5.10, [+0.34, +9.86]) above zero and the other four intervals containing zero.

Gains transfer across other models, benchmarks, and harnesses: zero-shot Mid-Harness improves Pass@3 and matches or improves Pass@1 across all seven settings, including Qwen3.5-9B and Nemotron3.5 Lightning on Terminus-2 and Nemotron3 Ultra on Terminal-Bench 2.1; on FeatureBench-Mini the TMAX-9B base agent succeeds in only 1.45% Pass@1 while zero-shot and distilled verification raise it to 5.80% and 7.25%, and on TMAX-27B distillation raises Pass@1 from 17.39% to 23.19% and Pass@3 from 26.09% to 39.13%. The work extends action verification from a single model family to generators from 4B to 550B, several benchmarks, and two harnesses, and reports domain and difficulty subgroup differences. Transfer experiments cover TerminalBench-Lite, Terminal-Bench 2.1, SWE-bench-Verified Mini, and FeatureBench-Mini; the authors also report that distillation improves TMAX-9B Pass@1 but lowers Pass@3 on FeatureBench-Mini and does not improve over zero-shot verification on Terminal-Bench 2.1.

Perspective

The result targets agents executing long-horizon tasks in terminal environments, under a deployment shape where the generator and harness stay fixed and sampling plus verification can be inserted before action execution; the authors evaluate in isolated benchmark environments and state that action verification should complement, not replace, safeguards such as restricted permissions, environment isolation, and human approval for sensitive or irreversible actions. Methodologically, Mid-Harness only changes the model-call wrapper without modifying model weights or serving architecture, so it stacks directly on existing parallel trajectory scaling (Best-of-N) and sequential trajectory scaling (Sequential Refine) without increasing their number of environment runs. The authors also note that distilled verifiers can be initialized from the corresponding generator backbone at 4B to 27B with the pairwise mechanism unchanged, offering a reproducible path for moving action verification to other terminal agents.

The evaluation lacks gold action labels, so verification correctness and candidate coverage can only be inferred indirectly through trajectory success under a strong verifier, which the authors list as a limitation. The offline diagnostics measure agreement with the teacher verifier rather than action correctness, and after distillation verification agreement remains lower in later states (54.07% at turns 17-32 versus 68.13% at turns 1-4), with remaining disagreements centered on command semantics and execution feasibility, which account for 67.4% of reviewed distilled-verifier failures. Decision-only (A/B-only) verification improves TMAX-9B Pass@1 at N=8 and lowers reference-priced cost, but is lower than its reasoning counterpart at 4B and 27B, and the authors note that missing usage records and separate live runs limit how precisely that comparison isolates the effect of response format. Four of the task-level bootstrap intervals contain zero, which the authors state does not establish an absence of improvement but means the analysis does not resolve a positive difference from zero; additionally, the distilled verifier improves Pass@1 but lowers Pass@3 on TMAX-9B FeatureBench-Mini, does not exceed zero-shot verification on Terminal-Bench 2.1, and reduces 27B Pass@1 by 10.0 points on scientific-computing tasks, indicating that gains are not uniform across domains.

Sources