Public articles linked to the same research event.
arXiv The authors introduce software-in-the-loop reconstruction (SWR), a self-supervised framework that draws reference outputs and verification targets from existing software workflows, instantiated with 500 workflows across 46 software families and six domains, where Qwen3.8-Max solves 838 tasks over three attempts and yields 1,422 verified trajectories oversampled into 3,000 reconstruction-only training examples, and supervised fine-tuning raises Qwen3.8-27B mean Terminal-Bench 2 performance from 47.94% to 53.56% while achieving the highest mean among four matched-token corpus controls on all four reported evaluations.
The authors introduce software-in-the-loop reconstruction (SWR), a self-supervised framework that draws reference outputs and verification targets from existing software workflows, instantiated with 500 workflows across 46 software families and six domains, where Qwen3.8-Max solves 838 tasks over three attempts and yields 1,422 verified trajectories oversampled into 3,000 reconstruction-only training examples, and supervised fine-tuning raises Qwen3.8-27B mean Terminal-Bench 2 performance from 47.94% to 53.56% while achieving the highest mean among four matched-token corpus controls on all four reported evaluations.
The authors introduce software-in-the-loop reconstruction (SWR), a self-supervised framework that draws reference outputs and verification targets from existing software workflows, instantiated with 500 workflows across 46 software families and six domains, where Qwen3.8-Max solves 838 tasks over three attempts and yields 1,422 verified trajectories oversampled into 3,000 reconstruction-only training examples, and supervised fine-tuning raises Qwen3.8-27B mean Terminal-Bench 2 performance from 47.94% to 53.56% while achieving the highest mean among four matched-token corpus controls on all four reported evaluations.
The authors introduce software-in-the-loop reconstruction (SWR), a self-supervised framework that draws reference outputs and verification targets from existing software workflows, instantiated with 500 workflows across 46 software families and six domains, where Qwen3.8-Max solves 838 tasks over three attempts and yields 1,422 verified trajectories oversampled into 3,000 reconstruction-only training examples, and supervised fine-tuning raises Qwen3.8-27B mean Terminal-Bench 2 performance from 47.94% to 53.56% while achieving the highest mean among four matched-token corpus controls on all four reported evaluations.