RIDE extrapolates RL teachers' hidden-state residuals and matches or beats the teacher on average across four base/teacher pairs
Synopsis
The work proposes RIDE (RL-Induced Direction Extrapolation): on student-generated trajectories it computes the layerwise hidden-state residual between an RL-trained teacher and its pre-RL checkpoint, then regresses the student's hidden states toward targets displaced beyond the teacher along that residual; across four base/RL-teacher pairs (R1-Distill-1.5B, Qwen3-4B, Llama-3.2-3B, Phi-4-mini) RIDE's mean Avg@16 on AIME24, AIME25 and AIMO approaches or exceeds the RL-trained teacher on every pair, the only compared method whose mean does so, and it consistently outperforms output-space extrapolation.
Interpretation
The paper identifies the RL-induced representation residual, the layerwise hidden-state difference between the teacher and its pre-RL checkpoint on identical on-policy contexts, as a measure of how RL changed the model's computation; in the same-initialization distillation setting both checkpoints are available. Prior work viewed OPD as KL-constrained RL with an implicit log-ratio reward and extrapolated in output or parameter space; here the RL-induced change is measured directly in representation space, at every layer, before the language-model head. The residual is computed deterministically from two forward passes (teacher and pre-RL checkpoint on the same student prefix) rather than sampled; the paper defines it at every layer and token position and describes its implementation.
RIDE displaces the representation-matching target beyond the teacher along this residual with a single coefficient, recovering OPRD exactly at coefficient 1; conditioned on a rollout, the regression is equivalent to maximizing a linear directional reward defined by the residual under a quadratic penalty centered at the teacher. Under a shared linear head, the output-space extrapolation target is the head-projected image of the RIDE target; RIDE keeps the reward-extrapolation view while replacing the sampled-token log-ratio with a deterministic hidden-state gradient. The paper derives Proposition 4.2: minimizing the squared distance to the target equals maximizing an inner-product reward minus a quadratic penalty, with the maximizer at the teacher plus a displacement along the residual; only the target differs, so coefficient 1 coincides with OPRD.
Across four base/RL-teacher pairs, RIDE's mean Avg@16 approaches or exceeds the RL-trained teacher on every pair, is the only method whose mean lies above the teacher line, and consistently outperforms OPRD and output-space extrapolation ExOPD. The paper separates where extrapolation happens from whether it happens: ExOPD and RIDE share the same coefficient and the same pre-RL reference and differ only in the space where the displacement is measured; ExOPD falls below its teacher on every pair and, on the three pairs where the teacher-to-base gap is small, below the untouched student as well. Table 1 reports Avg@16 on AIME24, AIME25 and AIMO for the four base/teacher pairs, with three independent training seeds per trained method averaged at run level; RIDE's margin over OPRD is a few points, and the paper notes that on three pairs the margin is within one across-seed standard deviation, with the Llama-3.2-3B pair resting mainly on AIMO.
Mechanistic analyses show that the language-model head attenuates the residual anisotropically and that output-space extrapolation amplifies sampled-token noise, whereas RIDE's gradient is deterministic given the rollout. Proposition 4.1 gives the conditional variance of the output-space advantage over the sampled token as growing quadratically with the extrapolation coefficient, while RIDE's per-position gradient is unchanged by resampling the next token; direction controls show the gain comes from the RL-induced residual itself. The residual retains only a small fraction of the head gain of an equal-norm isotropic direction, with its energy concentrated in the weakest head directions; random, reversed, mismatched-origin and trajectory-mismatched controls at fixed displacement magnitude all land near OPRD (within one point), while RIDE with the RL-induced residual is higher.
Perspective
The result targets the same-initialization distillation setting: the student starts from the teacher's pre-RL checkpoint, and teacher, base and student share architecture, tokenizer and representational origin, so hidden states can be compared layerwise; typical uses are distilling an expensive RL run back into its base model or merging domain experts that share a base. The method requires keeping the pre-RL checkpoint, uses a single global extrapolation coefficient, and performs representation regression at all layers and the last response positions; in implementation it adds one frozen forward pass on the teacher worker and transfers the same volume as OPRD. Evaluation is limited to mathematical reasoning benchmarks (AIME24, AIME25, AIMO) and one RL recipe, and the authors list relaxing these constraints and studying the safety and calibration of students trained on hidden-state targets as future work.
Open questions the paper itself raises include dependence on the pre-RL checkpoint and a shared representation space, the use of a single global coefficient, and evaluation limited to mathematical reasoning with one RL recipe. A careful reader will also watch that on three pairs the margin lies within one across-seed standard deviation, that the Llama-3.2-3B pair rests mainly on AIMO, and that the coefficient sweep uses a single seed per coefficient, differing from the three-seed means by about one standard deviation. In addition, the student is supervised on hidden-state targets that no model has actually produced, so its calibration and safety behavior are not guaranteed to match the teacher and the authors recommend independent evaluation before deployment; because the target is displaced beyond the teacher, biases or errors RL introduced could in principle be amplified rather than merely copied. This summary is based on the full paper text and the homepage evidence bundle and does not include training code or checkpoints (the authors state these will be released upon publication), so reproduction details still rest on the paper's appendices.
