Skip to main content
Back to timeline
arXivSource publication:

OASIS restricts self-distillation supervision to verified trajectories, widening its average gain over OPSD from 0.59 to 3.05 points across Qwen3-1.7B, 4B and 8B

Synopsis

Using truncated-continuation probes, the work finds that in on-policy self-distillation (OPSD) the teacher's advantage over the student is concentrated on failed trajectories (77% versus 37% recovery at 1.7B) and shrinks with scale, and therefore proposes OASIS: keeping the OPSD objective unchanged, it applies supervision only along the shortest outcome-verified on-policy trajectory and replaces the written reference solution with a distinct same-problem self-generated rollout, typically the student's own unverified attempt, improving mean Avg@12 over OPSD by 0.59, 1.86 and 3.05 points on Qwen3-1.7B, 4B and 8B across AIME 2024, AIME 2025 and HMMT 2025 while requiring no written solutions.

AI-generated editorial illustration: Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

Interpretation

The paper separates two roles that standard OPSD couples: the trajectory scaffold determines which prefixes receive supervision, while the teacher context determines the target distribution at those prefixes, and it measures the two separately with signal-replay and recovery probes. Earlier analyses of OPSD mostly approached the teacher side, for example decomposing the teacher signal into a reference-induced component and a question-conditioned component, or showing that a solution written for a different problem still provides supervision; this work instead asks where along the student's trajectory the supervision is applied. Inference-only probes on Qwen3-1.7B and Qwen3-4B with the frozen initial policy, 200 scaffolds per cell with 95% bootstrap intervals; the recovery probe cuts a rollout prefix and samples four continuations at temperature, counting a recovery when the last boxed answer matches ground truth.

The teacher's advantage is concentrated on failed trajectories: at 1.7B the gold-context teacher recovers a correct answer on 77% of failed prefixes versus 37% for the student alone, while empty and unrelated references give 37% and 32%; on verified prefixes the gap is small (99% versus 95%). This yields a measurable criterion: the value of the teacher signal depends on whether it can be reproduced from information available to the student, not on whether the reference solution is the standard one. Table 1 reports recovery rates by prefix type and context condition; on failed prefixes the final answer alone gives 61% and the written reasoning with the answer removed gives 51%, so both contribute and neither is available to the student at inference time.

OASIS keeps OPSD's loss, teacher, prompts, clipping constant, temperature and optimizer unchanged and changes only the distribution of supervised trajectories: it samples rollouts per problem, verifies their final answers, applies the loss along the shortest verified trajectory, and builds the teacher context from another rollout of the same problem, typically an unverified attempt. Unlike rejection-sampling fine-tuning, the verified trajectory is not the training target; the teacher distribution is. The method requires only a final answer per training problem, not a written solution. Across Qwen3-1.7B, 4B and 8B, OASIS improves over OPSD by 0.59, 1.86 and 3.05 points on average and over the base model by 3.64, 3.84 and 3.19 points, while OPSD's gain over the base model falls from 3.05 to 1.98 to 0.14; the authors state that the 0.59-point gap at 1.7B is within the range sampling variation can produce on 30 problems and is therefore not conclusive on its own.

A training-budget comparison shows that at 4B and 8B, OASIS matches or exceeds OPSD and AVSD on average with roughly a quarter of the supervised problems and no written solutions, at the price of substantially more generation. This trades OPSD's dependence on written solutions for a dependence on verifiable answers plus more sampling, which extends the method to datasets that provide only checkable answers. In Table 4, OPSD and AVSD each see 12.8K problems and train on all of them, while OASIS sees 6.4K problems, trains on about 3.2K and generates 51.2K rollouts; at 1.7B OASIS generates 43.5M tokens and supervises 1.98M, against 2.76M generated and 2.76M supervised for OPSD, with verified scaffolds averaging 603 versus 859 tokens.

Perspective

The result targets settings where a mathematical reasoning model is trained and each training problem has only a verifiable final answer rather than a written solution; it is validated on Qwen3-1.7B, 4B and 8B and on AIME 2024, AIME 2025 and HMMT 2025, and the authors name extension to settings with inexpensive outcome verification but no written solutions, such as code with unit tests, as a natural next step. For a practitioner this means sampling plus verification can replace written-solution annotation, at the cost of higher generation: OASIS samples eight rollouts per problem, generating 43.5M tokens over 100 steps at 1.7B against 2.76M for OPSD, and the authors note that an ablation in the appendix halves this cost where generation is the binding constraint.

How the shrinking teacher advantage behaves at larger scales is only shown up to 8B; the 0.59-point OASIS-over-OPSD gap at 1.7B falls within sampling variation, which the authors themselves say is not conclusive on its own; 4B and 8B are single training runs, and although Table 3 uses three evaluation seeds to show evaluation variability is small relative to the OASIS-OPSD gap, training-level variability awaits more repeats. The clipping analysis shows that per-vocabulary-entry clipping removes the teacher's upward pull on a clipped token and reverses the sign of its logit gradient, which the authors offer as one possible explanation for late-training drift and suggest a larger threshold or clipping per-token divergence instead; whether that change alters the conclusions remains open. The verifier is also imperfect, so an unverified rollout is more precisely a trajectory not verified as correct rather than a necessarily incorrect solution.

Sources