SFT can generalize too: locking updates to one on-policy direction halves training time and edges past GRPO
Lead
Constraining supervised fine-tuning updates to the direction identified by on-policy training lets SFT match GRPO's generalization while cutting training time by more than half.
Story
The generalization gap of supervised fine-tuning (SFT, which optimizes a model on fixed ground-truth responses) can be closed by a direction constraint alone, without changing its training objective. SFT was previously held to memorize while reinforcement learning generalizes, and efforts to improve it focused on data selection, weighted losses, or adding negative samples. Across Qwen3-1.7B/4B/8B and DeepSeek-R1-Distill-Llama-3-8B, direction-constrained OPSFT consistently beats vanilla SFT on math and code benchmarks and reaches the level of the GRPO run that supplies the direction.
The recipe runs a few on-policy steps, records the sign of each parameter's cumulative update, then keeps only the SFT gradient components whose signs match. Earlier work used only the update locations of on-policy training, that is, which parameters change, and treated them as a byproduct. Constraining only the update locations performs worse than vanilla SFT, while extending the constraint from location to direction yields large generalization gains, showing the direction rather than the location carries the effect.
The direction constraint also lets SFT keep absorbing newly acquired high-quality trajectories without disrupting what on-policy training already learned. Applying SFT directly to a post-trained model degrades it, for example dropping Qwen3-4B mean accuracy from 38.96 to 36.25. Continuing training along the original on-policy direction raises the same model to 41.36 mean accuracy, with consistent results on code tasks.
What to watch
A researcher wanting to reuse this can run a few on-policy steps in the target domain, extract the cumulative update signs, and use them to constrain SFT, obtaining near-on-policy generalization on verifiable tasks such as math or code at shorter training time. Engineering teams running post-training pipelines can attach the same direction constraint to existing SFT code to keep absorbing newly collected high-quality trajectories, instead of restarting the whole SFT-then-reinforcement-learning process from the base model.
How much the direction constraint depends on trajectory quality remains open: the experiments use reasoning trajectories generated by a teacher model substantially stronger than the model being trained, whereas on-policy paradigms in practice train on trajectories from the current policy, which may differ in diversity, correctness, and suitability. Existing studies measuring reasoning-trajectory quality and suitability are limited, so how trajectory properties affect the identified direction and OPSFT performance is worth watching.
