RLVR post-training is proved to permit unbounded language drift while supervised fine-tuning does not, with drift appearing only on novel reasoning tasks the base model cannot perform
Related research and updatesSynopsis
This work characterizes when language drift arises in RLVR post-training: it proves theoretically that RLVR optimization pressure permits unbounded language drift whereas supervised fine-tuning does not, shows empirically that drift arises specifically during RLVR on novel reasoning tasks where the target behavior cannot be drawn out of the base model, and further proves that language drift cannot be constrained without constraining expected reward.
Figure 1 : SFT (left) on prompt/CoT/answer triples ( x , z , y ) (x,z,y) does not permit a model π ( SFT ) \pi^{(\textit{SFT})} to drift arbitrarily far from a human language distribution P ( HL ) {P^{(\textit{HL})}} while minimizing cross-entropy with a target distribution D ( SFT ) {{D}^{(\textit{SFT})}} . RLVR (right) places no optimization pressure on the CoT z z , allowing arbitrary language drift away from natural language in the reasoning trace. The example completions in this figure were drawn from Llama-3.2-1B models trained via SFT and RLVR (respectively) on GSM8K.
arXivInterpretation
The paper proves theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-tuning does not. Language drift was documented but its causes were "poorly understood"; this work moves from describing the phenomenon to a formal distinction at the level of the optimization objective, showing RLVR and supervised fine-tuning differ in whether unbounded drift is permitted. The evidence is a theoretical proof, stated in the abstract as "we prove theoretically"; proof details, assumptions, and theorem counts are not given.
Language drift arises specifically during RLVR on novel reasoning tasks, that is, when the target behavior cannot be drawn out of the base model. It ties the occurrence of drift to task novelty and the base model's capability boundary, rather than treating drift as a general side effect of training. The evidence is empirical, stated in the abstract as "we show empirically"; model scale, task sets, and drift metrics are not provided.
The paper proves it is not possible to constrain language drift without constraining expected reward, so improving chain-of-thought monitorability during frontier RLVR post-training necessarily harms performance. It elevates the monitorability-performance trade-off from an empirical observation to a provable impossibility result, giving formal support to the existing concern that drift can impair CoT monitorability. The evidence is a theoretical proof, stated in the abstract as "we prove that it is not possible"; no quantitative form or bound for the trade-off is given.
Perspective
The result targets reasoning models post-trained with reinforcement learning from verifiable rewards, especially novel reasoning tasks where the target behavior cannot be drawn directly out of the base model; for such settings it indicates a structural trade-off between monitorability and expected reward, relevant to researchers and engineering teams designing post-training objectives and chain-of-thought monitoring pipelines.
The abstract does not state the assumptions behind the theoretical proofs, the reward-model form, or the task distribution, nor does it give model scale, task sets, or drift metrics for the empirical part; the precise boundaries of "unbounded drift" and "cannot be constrained," and the reward and task conditions under which the conclusions hold, still need confirmation in the full text.
