ROSS reuses discarded historical self-generated rollouts, lifting Qwen3.6-35B-A3B's six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%
Synopsis
ROSS introduces a selective-supervision relearning procedure that takes historical self-generated rollouts saved from domain-specific RL, multi-teacher on-policy distillation (MOPD), or agentic RL, uses an outcome verifier to keep successful trajectories and an LLM reviewer to mark which model-generated spans are worth imitating, then applies loss only to those selected tokens while keeping the full trajectory as context, improving upstream checkpoints through an offline SFT stage without new policy rollouts and raising Qwen3.6-35B-A3B's six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%.
Interpretation
The paper establishes historical self-generated rollouts as a reusable source of training experience rather than a stale artifact to discard once the policy advances. Prior methods such as RFT, RAFT, ReST-MCTS*, RLoop, and RIFT largely treat a rollout as an indivisible supervision unit, selecting, refining, or reweighting trajectories by overall outcome or reward; ROSS instead separates the question into two levels: which historical rollouts are worth revisiting, and which segments within them are worth learning. Two preliminary analyses are reported: on 100 math problems, historical rollouts have substantially lower token-level NLL under the current policy than external-teacher GLM-5.2 responses and remain close to current-policy rollouts; on the hardest training problems with at least one historical success, the current policy still solves them but with a low average success rate, indicating behaviors that remain within its behavioral support yet are not reliably expressed.
ROSS uses a masked teacher-forcing objective with full context and selective targets, computing loss only on reviewed model-generated spans. Unlike Positive-Rollout SFT, which supervises all eligible tokens in verifier-positive trajectories, ROSS keeps the original prefix (including earlier mistakes, abandoned attempts, and environment feedback) but removes confirmed errors and redundant exploration from the loss; single-turn responses use positive-span extraction while agentic histories use defect localization, both compiled into the same token-level mask. The method combines trajectory-level selection by an outcome verifier with segment-level annotation by an LLM reviewer, followed by deterministic checks for source consistency and token alignment; candidates are rejected when no valid target can be recovered without retaining a confirmed error or introducing missing reasoning.
Across domain-specific RL, MOPD, and agentic RL, ROSS outperforms upstream checkpoints and relearning baselines, and the gains decompose into trajectory filtering and token-level masking. Three sequential contrasts (replay, filtering, masking) separate the effects: ROSS w/o mask, which only filters trajectories, already raises the MOPD average from 57.70 to 58.49, slightly above upstream; with the replay set fixed, full ROSS improves further over ROSS w/o mask by 3.71 points on MOPD, 1.91 on LCB Gen, 0.43 on OJBench, and 0.66 on Math. Under domain-specific RL, Math average rises from 75.19 to 76.98 and Code average from 42.09 to 44.21; under MOPD the six-benchmark average rises from 58.40 to 62.20; in the agentic setting SWE-bench Verified rises from 64.20 to 68.40, a 4.20-point gain over the upstream checkpoint.
The value of historical trajectories does not depend on inheriting upstream parameter updates, and excluded supervision does not vanish as RL progresses. Applying identical historical trajectories, supervision masks, and SFT configurations from either Base or Upstream, Base-initialized ROSS performs comparably to Upstream-initialized ROSS across evaluated settings, suggesting the experience transfers across initializations; meanwhile the semantic composition of excluded content differs by domain, with corrected mistakes dominating both early and late Math samples and redundant exploration becoming more prevalent from early to late in Code. The token-weighted target exclusion rate (TER) decreases only modestly from early to late in Math and increases in Code; in the separate 4B Math run in Appendix H, the upstream checkpoint exhibits repeated answer emission, and ROSS reduces the repetition rate from 84.5% to 21.5% and the mean boxed-answer count from 628.3 to 63.0 while raising Avg. Math from 52.05 to 63.65.
Perspective
This work targets post-training pipelines that have completed a round of self-rollout training and saved both checkpoints and rollouts, and it applies to tasks with an available outcome verifier such as math, code generation, instruction following, and software engineering; ROSS runs as an additional offline SFT stage after domain-specific RL, MOPD, or agentic RL, requires an LLM reviewer for span annotation, and relies on deterministic checks for source and token alignment. For teams that want to extract further value from existing training experience without resampling rollouts, or to transfer behaviors discovered in one run to a different initialization, this offers an actionable path; the agentic setting additionally requires retaining full interaction histories and a longer context budget.
Span-level annotation depends on the LLM reviewer's judgment; the paper constrains it with independent audits and deterministic boundary checks, but the reviewer's stability across task distributions remains a question readers will watch. The categorization of excluded content is based on sampled trajectories covering two windows of the observed run rather than the full training process. In cross-domain results, Code ROSS decreases AIME 2025 while Math ROSS trails full-replay Positive-Rollout SFT on cross-domain benchmarks, suggesting the distribution of gains varies with setting. In addition, the suppression of repetitive output in the 4B Math run comes from a separate appendix experiment, and its generalization to larger models and other domains would need more evidence.
