SWR self-supervises training environments from 500 scientific software workflows, lifting Qwen3.8-27B on Terminal-Bench 2 from 47.94% to 53.56%
Related research and updatesSynopsis
The authors introduce software-in-the-loop reconstruction (SWR), a self-supervised framework that draws reference outputs and verification targets from existing software workflows, instantiated with 500 workflows across 46 software families and six domains, where Qwen3.8-Max solves 838 tasks over three attempts and yields 1,422 verified trajectories oversampled into 3,000 reconstruction-only training examples, and supervised fine-tuning raises Qwen3.8-27B mean Terminal-Bench 2 performance from 47.94% to 53.56% while achieving the highest mean among four matched-token corpus controls on all four reported evaluations.
Figure 2 : Software-in-the-loop reconstruction. Public observations are the input–output examples released to the agent. They guide construction of a reusable program. Withheld inputs supply the reference outputs used for verification.
arXivInterpretation
Introduces software-in-the-loop reconstruction, which treats existing executable software workflows as the source of reference behavior, executes multiple input configurations per workflow, and partitions cases into public observations and hidden evaluations. Previously, building training environments for terminal agents required authoring executable reference behavior and a domain-specific verifier for each task, which meant repeated engineering and limited reuse; SWR instead obtains reference outputs and verification targets from existing software workflows. The abstract describes the full pipeline: given the instruction, input schema, and public input-output observations, an agent constructs an editable program without access to the source workflow, and the candidate is evaluated on hidden configurations against workflow outputs.
Designs a hierarchical verifier that combines domain-specific semantic comparison, structural validity, and anti-shortcut checks, with public feedback supporting iterative revision. Verification no longer judges only whether an artifact looks plausible; it distinguishes semantic correctness from superficially plausible artifacts and constrains candidate programs through anti-shortcut checks. The abstract names the three verifier components and the iterative role of public feedback, but does not report component weights or ablations in the instantiation.
Instantiates SWR with 500 workflows and 46 software families across six domains, where Qwen3.8-Max solves 838 tasks across three attempts per task, produces 1,422 verified trajectories, and oversamples them into 3,000 reconstruction-only training examples. The construction admits additional workflows and configurations without authoring a reference solution for each task, so the supervision signal can scale with existing software. The abstract reports concrete counts: workflows, software families, domains, solved tasks, verified trajectories, and training examples.
Supervised fine-tuning of Qwen3.8-27B on these examples improves mean Terminal-Bench 2 performance from 47.94% to 53.56% across three seeds and achieves the highest mean among four matched-token corpus controls on all four reported evaluations. The results indicate that existing scientific software can serve as scalable, behaviorally verified supervision for terminal agents rather than requiring hand-authored per-task environments. The abstract reports the three-seed mean change and the control comparison across four evaluations, but does not list per-evaluation values, variance, or statistical tests.
Perspective
The work targets settings where terminal agents must perform tasks in science and other specialized domains, and it applies where executable software workflows exist whose inputs and outputs can be recorded in structured form. Beneficiaries include engineering teams building training environments for terminal agents and domain developers who want to reuse existing scientific software instead of authoring per-task reference solutions. The construction described in the abstract admits additional workflows and configurations without authoring a reference solution for each task, so its scaling path is to bring more existing software into the same self-supervised pipeline. The reported fine-tuning gains pertain to Qwen3.8-27B on Terminal-Bench 2 and the four reported evaluations, and are results within that setting.
The abstract does not list per-evaluation values, variance, or statistical tests for the four reported evaluations, nor does it state the weights or ablations of the three verifier components, so the relative contribution of each verification stage remains an open question. The composition of the six domains and 46 software families, the split ratio between public observations and hidden evaluations, and the shortcut types covered by the anti-shortcut checks are not detailed in the abstract. In addition, this reading is limited to the abstract and does not include the body, figures, or appendix, so these details cannot be verified from the available text; readers assessing transferability should consult the corresponding parts of the original.
