Skip to main content
Back to timeline
arXivSource publication:

CoTrace matches training trajectories to the runtime, lifting Qwen3.5-9B on Tmax from 78 to 88 solved tasks (RL variant reaches 90)

Synopsis

The work introduces CoTrace, a harness-aware data recipe for alternating model–harness co-evolution that explicitly governs trajectory routing, provenance matching, and curriculum refresh: execution failures are routed to harness synthesis while policy supervision uses only verified successes matched to the adopted runtime. On the 102-task Tmax promotion split, Qwen3.5-9B advances from 78 to 88 solved tasks under supervised fine-tuning and reaches 90 with an online reinforcement variant, whereas much larger corpora pooled across sibling harnesses yield no accepted model update.

Source-provided article image: CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution
Figure 1 ·

Figure 1: Overview of CoTrace and the shared data interface in model–harness co-evolution. Bottom: the closed loop. A task agent with the adopted harness H ∗ H^{*} and policy weights θ \theta solves executable terminal tasks and fills a shared history of trajectories; harness evolution reads that history to edit prompts, processors and tools and selects the next H ∗ H^{*} , while model training updates θ \theta before the subsequent round of search. Top: the data recipes that turn the history into model-training data, from the baseline supervised fine-tuning (SFT) recipes that pool every successful search trajectory (all-evolve) or successes across sibling candidate harnesses (mixed siblings), to CoTrace-SFT, which keeps only trajectories whose provenance matches H ∗ H^{*} and tops them up with fresh rollouts under H ∗ H^{*} , and CoTrace-RL, which learns online by reinforcement learning (RL) from rewards on frontier tasks rolled out under H ∗ H^{*} .

arXiv

Interpretation

CoTrace explicitly splits execution data during co-evolution: failure traces are clustered by earliest unrecovered failure and routed to harness synthesis, while policy supervision uses only verified successes whose fingerprint matches the adopted runtime, topped up with fresh rollouts under that same harness when coverage is insufficient. Prior harness–model co-evolution approaches treat trajectories produced during harness search as an undifferentiated replay buffer; CoTrace makes trajectory provenance (a fingerprint hashing prompt template, tool bindings, and processors) a first-class variable in data selection. In a controlled tournament-search comparison, the mixed-siblings corpus supplies 149–308 trajectories per iteration yet yields 0 accepted model updates, whereas the matched recipe uses only 30–50 and contributes 2 accepted updates at 9B.

Component-wise promotion (Ratchet) evaluates each candidate on the frozen 102-task split with the other component held fixed, so every accepted gain is attributable to a single component and regressive updates are rejected. The protocol decouples harness search and policy training into auditable alternating channels with a recorded promotion decision per update, rather than reporting only a pair-level total. Across 17 harness candidates from 8 chains that had already won their evolve-set search, only 9 improved the promotion split, with a mean effect of +0.18 tasks; a cross-evaluation shows model and harness gains adding rather than interacting in one completed transition.

Curriculum refresh (Refresh) retires a task only once it is both solved and harvested into the training corpus, keeping the evolve set concentrated on the moving residual frontier. Tying retirement to the availability of usable training signal prevents tasks from leaving before their success becomes trainable data while still requiring evidence of mastery. In one chain the evolve-set score fell 34→26→17 of 50 while the promotion score rose 75→81→83 of 102; the reinforcement chain's third iteration retired 40 of 50 tasks at once, and its final candidate is the only stage that gives tasks back (85 against the incumbent 90).

Cross-domain transfer is governed by the co-evolved model–harness pair rather than the policy in isolation: the same checkpoint performs materially differently under different runtimes. The transfer analysis moves from comparing checkpoints alone to comparing model–harness pairs, showing that checkpoint-only gains largely disappear under a foreign runtime. On SWE-bench Lite the three checkpoints resolve 34.7%, 24.7%, and 36.7% under mini-swe-agent; pairing each with its adopted harness improves every result, most strongly for SFT (24.7% to 35.0%), while reducing its no-patch outcomes from 155 to 64, and the reinforcement pair reaches 41.0%.

Perspective

The recipe targets agent post-training on executable terminal tasks with binary verifier rewards, in settings where model and runtime must be updated alternately. It lets researchers and engineers route failure evidence to harness synthesis, restrict policy supervision to verified successes matched to the deployed runtime, and advance the curriculum only after a task is both solved and harvested. At 9B scale, a compact matched corpus yields model-side gains at 47 GPU-hours per iteration; at 4B scale, when clean successful trajectories are sparse, online reinforcement still obtains a relative reward signal from the same policy. Cross-domain evaluations indicate that deploying a checkpoint within its co-evolved runtime suppresses early execution faults and no-patch outcomes, a finding scoped to the runtimes and benchmarks studied.

Evidence comes from a small number of co-evolution chains each run once; the promotion split has 102 tasks and single evaluations vary by roughly two tasks, so one- or two-task differences between chains should not be read as rankings. The supervised and reinforcement recipes were not matched in compute, and harness search depends on a proprietary meta-agent whose proposals are released only as accepted configurations. The study covers one model family at two scales, one task source, and two external benchmarks, so transfer findings describe these runtimes rather than runtimes in general. Chains stop as the curriculum saturates, leaving open whether a larger task reservoir would extend the gains. Separately, the reinforcement stage initially exhibited a data-path defect in which the policy never emitted the terminal submission action and every reward was zero; reward became non-zero only after the prompt schema and sandbox working directory were rebuilt, a reminder that data-plane issues can masquerade as capability bottlenecks.

Sources