Full text not available for this paper
Summary (Overview)
- Problem: On-policy distillation (OPD) aims to transfer capabilities from a reinforcement learning (RL)-trained teacher to a student model sharing the same initialization. While generalized variants allow students to surpass the teacher via output-space extrapolation, this approach suffers from anisotropic attenuation of representation changes by the language-model head and noise amplification from sampled-token log-probability ratios.
- Key Insight: RL shifts a model's internal representations relative to its base checkpoint, and this shift's direction can be measured at every layer. This direction, not just the destination, can be extrapolated.
- Proposed Method: RIDE (RL-Induced Direction Extrapolation) — extrapolates the RL-induced change directly in representation space. At every layer and token position, RIDE computes the residual between the teacher and its pre-RL checkpoint and regresses the student's hidden states toward targets displaced beyond the teacher along this residual.
- Main Results: Across four base/RL-teacher pairs (spanning different scales, architectures, and pre-training lineages), RIDE approaches or exceeds the RL-trained teacher on every pair, consistently outperforming output-space extrapolation (ExOPD) and representation-matching (OPRD) baselines.
- Theoretical Contribution: RIDE's objective is equivalent to maximizing a linear directional reward under a quadratic penalty centered at the teacher, and its gradient is deterministic given a rollout, unlike the noisy sampled-token advantage of output-space extrapolation.
Introduction and Theoretical Foundation
Background and Motivation
- On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories, reducing exposure bias and providing dense learning signals. It has become a standard post-training stage for language models.
- A common use case is transferring the gains of an RL-trained teacher back to a student that shares its initialization (e.g., merging domain experts or distilling an expensive RL run).
- Beyond-the-teacher learning: Standard OPD treats the teacher as the endpoint. However, the teacher-to-reference log-probability ratio acts as an implicit reward; up-weighting it can yield students that outperform the teacher. Because RL updates are small, structured, and consistent across runs, an RL-trained teacher supplies a direction as well as a destination.
The Problem with Output-Space Extrapolation
- Head attenuation: The language-model head reweights representation changes according to its anisotropic singular spectrum. The 512 weakest directions of the head carry 79.8% of the teacher-to-base residual's hidden-state energy but only 30.0% of its centered-logit energy. Thus, output-space objectives supervise much of the residual only weakly and leave earlier layers unconstrained.
- Noise amplification: The sampled-token log-ratio is length-biased and prone to overoptimization. Scaling it by a global coefficient amplifies extreme rewards and destabilizes training.
Theoretical Foundation
- Same-initialization setting: The teacher is obtained by RL from a base checkpoint (), and the student is initialized at that checkpoint (). All three models share architecture, tokenizer, and representational origin.
- RL-induced residual: On a common prefix, the difference measures what RL changed on that context — in output space it is the implicit reward ; in representation space it is a vector at every layer, taken before the head.
Methodology
The RIDE Objective
The core idea is to displace the training target from the teacher along the RL-induced change:
- At , the target is the teacher (recovering standard OPD/OPRD).
- For , the target lies beyond the teacher, continuing the change RL initiated.
Applying this to hidden states yields the RIDE target:
where is the RL-induced residual at layer and position .
Training objective (representation-space regression):
where and stops gradients through the target. All layers and the final response positions are supervised.
Why Hidden States?
-
Head attenuation: The output-space target is the head-projected image of the RIDE target: . However, a logit target constrains hidden states only weakly along directions the head amplifies least, and imposes no constraint on layers below the final one.
-
Noise properties: Output-space extrapolation uses a sampled advantage where and .
Proposition 4.1 (Conditional variance under extrapolation):
The second term does not vanish as unless is constant on the support of . In contrast, the RIDE gradient is deterministic given a rollout, so its conditional variance is zero.
Optimization Interpretation
Proposition 4.2 (Reward-and-penalty form): Minimizing is equivalent to maximizing:
where . The reward is the inner product of the student's displacement from the teacher with the RL-induced residual, and the penalty is quadratic in the same displacement. The maximizer is .
Empirical Validation / Results
Experimental Setup
- Models: Four base/RL-teacher pairs: R1-Distill-1.5B → JustRL-1.5B, Qwen3-4B → Just-Qwen3-4B, Llama-3.2-3B → Just-Llama-3.2-3B, Phi-4-mini → Just-Phi-4-mini.
- Training: Prompts from DAPO-Math-17K, 500 training steps, 3 independent seeds (14, 42, 2027).
- Evaluation: Avg@16 on AIME24, AIME25, and AIMO (AMC 2022–2023).
- Baselines: OPD (top-1 and top-16), OPRD (representation matching, ), ExOPD (output-space extrapolation, ).
Main Results (Table 1)
| Method | R1-Distill-1.5B | Qwen3-4B | Llama-3.2-3B | Phi-4-mini |
|---|---|---|---|---|
| Teacher | 55.30 | 65.59 | 13.01 | 17.96 |
| Student (untouched) | 39.00 | 62.82 | 6.99 | 15.13 |
| OPD top-1 | 52.90 | 59.24 | 1.08 | 12.90 |
| OPD top-16 | 52.53 | 59.74 | 5.97 | 16.76 |
| OPRD | 54.50 | 62.01 | 10.94 | 17.33 |
| ExOPD | 49.87 | 48.75 | 6.95 | 12.22 |
| RIDE | 56.38 | 66.07 | 13.33 | 18.30 |
Key findings:
- RIDE is the only method whose mean lies above the teacher line on all four pairs.
- RIDE outperforms OPRD by 0.97–4.06 points, isolating the benefit of residual extrapolation.
- ExOPD falls below its teacher on every pair and below the untouched student on three pairs (by 14.1 points on Qwen3-4B), confirming the noise amplification predicted by Proposition 4.1.
Coefficient Sweep (Figure 4)
- RIDE: Final Avg@16 rises with from 48.2 at to 54.3 at , peaking at 55.4 at . Every improves over OPRD. Degradation beyond is graceful ( finishes at 52.4).
- ExOPD: Best at (52.9), harmed by every . The run peaks early then declines to 49.9; collapses to 46.2 with format score falling from 92% to 64%.
Mechanistic Analysis (Figure 5)
- Head attenuation: The residual retains only 0.59 of the head gain of an equal-norm isotropic direction. The 512 weakest head directions hold 79.8% of the residual's hidden-state energy (vs. 33.3% under isotropy). These directions receive 73.0% of the representation-loss gradient but only 48.6% of the output-KL gradient.
- Student alignment: The student's update aligns with the residual (cosine 0.954), with a projection onto the residual of 1.69, above the target coefficient of 1.25 — the student continues past the target along the same direction.
- Conditional variance: The ExOPD conditional variance rises 12.6× from to , with the extrapolation term dominating beyond . The RIDE per-position gradient norm is invariant to next-token resampling (variance = 0).
Direction Ablation (Table 3)
Holding displacement magnitude fixed and altering only direction:
- Random direction: 55.04 (vs. OPRD 54.50)
- Reversed direction (): 53.90
- Mismatched origin (different base model): 54.75
- Trajectory-mismatched residual: 55.12
- RIDE (RL-induced residual): 56.38
The RL-induced residual itself is responsible for RIDE's gain.
Theoretical and Practical Implications
Theoretical Significance
-
Representation-space extrapolation as reward maximization: RIDE formalizes the notion that RL-induced representation residuals encode a usable "direction" for improvement, providing a representation-space counterpart to output-space reward extrapolation with a deterministic gradient.
-
Diagnosis of output-space extrapolation failure: Proposition 4.1 provides a rigorous explanation for why output-space extrapolation destabilizes training: the sampled-advantage noise scales with regardless of signal strength, and the damage is largest when the teacher-to-base gap is small.
-
Head attenuation quantification: The paper provides concrete measurements showing how the language-model head anisotropically attenuates representation changes, with the residual concentrating in the weakest head directions.
Practical Implications
- RIDE requires only the pre-RL checkpoint and a shared representation space, making it a practical drop-in replacement for OPRD pipelines with a single additional frozen forward pass.
- The method is stable across a range of extrapolation coefficients (), unlike output-space extrapolation which degrades sharply.
- The transfer volume to the trainer equals that of OPRD (a single hidden-state tensor), with no additional communication overhead.
Conclusion
Main Takeaways
- An RL-trained teacher supplies a direction as well as a destination; the layerwise hidden-state difference between teacher and pre-RL checkpoint on identical prefixes measures what RL changed.
- RIDE displaces the representation-matching target beyond the teacher along this residual with a single coefficient , recovering OPRD exactly at .
- Measuring the displacement before the head matters because the head attenuates the residual anisotropically and leaves earlier layers unconstrained.
- RIDE approaches or exceeds its RL-trained teacher on all four base/teacher pairs, consistently outperforms OPRD and output-space extrapolation, and remains stable where the latter collapses.
Limitations and Future Work
- RIDE requires the pre-RL checkpoint and a shared representation space.
- Uses a single global (no per-layer or per-position adaptation).
- Evaluated only on mathematical reasoning with one RL recipe.
- Future work: relaxing these constraints, studying the safety and calibration of students trained on hidden-state targets, and exploring per-layer or adaptive extrapolation coefficients.
Related papers
- Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Domain-normalized multi-teacher distillation, which rescales teacher feedback by log-ratio dispersion, outperforms label-routed MOPD by up to 3.08 points across model scales.
- Reward Hacking Challenges Oversight of Autonomous Research Agents
Autonomous research agents reward-hack 30.5% of research-pipeline tasks spontaneously, and detailed reviewer feedback doubles adaptive evasion rates to 40.5%.
- Post-Training Leaves Behavioral Shadows on Unrelated Decisions
Post-training updates leave behavioral shadows on unrelated inputs that can be extracted via single-word queries to transfer capabilities, yielding +5.34 points on HumanEval+.