Full text not available for this paper

Summary (Overview)

  • Problem: On-policy distillation (OPD) aims to transfer capabilities from a reinforcement learning (RL)-trained teacher to a student model sharing the same initialization. While generalized variants allow students to surpass the teacher via output-space extrapolation, this approach suffers from anisotropic attenuation of representation changes by the language-model head and noise amplification from sampled-token log-probability ratios.
  • Key Insight: RL shifts a model's internal representations relative to its base checkpoint, and this shift's direction can be measured at every layer. This direction, not just the destination, can be extrapolated.
  • Proposed Method: RIDE (RL-Induced Direction Extrapolation) — extrapolates the RL-induced change directly in representation space. At every layer and token position, RIDE computes the residual between the teacher and its pre-RL checkpoint and regresses the student's hidden states toward targets displaced beyond the teacher along this residual.
  • Main Results: Across four base/RL-teacher pairs (spanning different scales, architectures, and pre-training lineages), RIDE approaches or exceeds the RL-trained teacher on every pair, consistently outperforming output-space extrapolation (ExOPD) and representation-matching (OPRD) baselines.
  • Theoretical Contribution: RIDE's objective is equivalent to maximizing a linear directional reward under a quadratic penalty centered at the teacher, and its gradient is deterministic given a rollout, unlike the noisy sampled-token advantage of output-space extrapolation.

Introduction and Theoretical Foundation

Background and Motivation

  • On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories, reducing exposure bias and providing dense learning signals. It has become a standard post-training stage for language models.
  • A common use case is transferring the gains of an RL-trained teacher back to a student that shares its initialization (e.g., merging domain experts or distilling an expensive RL run).
  • Beyond-the-teacher learning: Standard OPD treats the teacher as the endpoint. However, the teacher-to-reference log-probability ratio acts as an implicit reward; up-weighting it can yield students that outperform the teacher. Because RL updates are small, structured, and consistent across runs, an RL-trained teacher supplies a direction as well as a destination.

The Problem with Output-Space Extrapolation

  • Head attenuation: The language-model head reweights representation changes according to its anisotropic singular spectrum. The 512 weakest directions of the head carry 79.8% of the teacher-to-base residual's hidden-state energy but only 30.0% of its centered-logit energy. Thus, output-space objectives supervise much of the residual only weakly and leave earlier layers unconstrained.
  • Noise amplification: The sampled-token log-ratio is length-biased and prone to overoptimization. Scaling it by a global coefficient amplifies extreme rewards and destabilizes training.

Theoretical Foundation

  • Same-initialization setting: The teacher is obtained by RL from a base checkpoint (πT=RL(πB)\pi_T = \text{RL}(\pi_B)), and the student is initialized at that checkpoint (πθ0=πB\pi_{\theta_0} = \pi_B). All three models share architecture, tokenizer, and representational origin.
  • RL-induced residual: On a common prefix, the difference sT,t−sB,ts_{T,t} - s_{B,t} measures what RL changed on that context — in output space it is the implicit reward log⁡πT−log⁡πB\log \pi_T - \log \pi_B; in representation space it is a vector at every layer, taken before the head.

Methodology

The RIDE Objective

The core idea is to displace the training target from the teacher along the RL-induced change:

τt(λ)=sT,t+(λ−1)(sT,t−sB,t)=λsT,t+(1−λ)sB,t,λ≥0\tau_t(\lambda) = s_{T,t} + (\lambda - 1)\left(s_{T,t} - s_{B,t}\right) = \lambda s_{T,t} + (1 - \lambda) s_{B,t}, \quad \lambda \geq 0
  • At λ=1\lambda = 1, the target is the teacher (recovering standard OPD/OPRD).
  • For λ>1\lambda > 1, the target lies beyond the teacher, continuing the change RL initiated.

Applying this to hidden states yields the RIDE target:

ht(l)⋆=hT,t(l)+(λ−1)Δt(l)=λhT,t(l)+(1−λ)hB,t(l)h^{(l)\star}_{t} = h^{(l)}_{T,t} + (\lambda - 1)\Delta^{(l)}_{t} = \lambda h^{(l)}_{T,t} + (1 - \lambda) h^{(l)}_{B,t}

where Δt(l)=hT,t(l)−hB,t(l)\Delta^{(l)}_{t} = h^{(l)}_{T,t} - h^{(l)}_{B,t} is the RL-induced residual at layer ll and position tt.

Training objective (representation-space regression):

Lλ(θ)=Ex,y^∼πθ[1∣Llayer∣∑l∈Llayer1M∑t=1Tmt1d∥hθ,t(l)−sg(ht(l)⋆)∥22]\mathcal{L}_{\lambda}(\theta) = \mathbb{E}_{x, \hat{y} \sim \pi_{\theta}}\left[\frac{1}{|\mathcal{L}_{\text{layer}}|} \sum_{l \in \mathcal{L}_{\text{layer}}} \frac{1}{M} \sum_{t=1}^{T} m_t \frac{1}{d} \left\| h^{(l)}_{\theta,t} - \text{sg}\left(h^{(l)\star}_{t}\right) \right\|_2^2\right]

where M=∑tmtM = \sum_t m_t and sg\text{sg} stops gradients through the target. All LL layers and the final kk response positions are supervised.

Why Hidden States?

  1. Head attenuation: The output-space target is the head-projected image of the RIDE target: Whead(h~T+(λ−1)Δ~)=zT+(λ−1)(zT−zB)W^{\text{head}}(\tilde{h}_T + (\lambda-1)\tilde{\Delta}) = z_T + (\lambda-1)(z_T - z_B). However, a logit target constrains hidden states only weakly along directions the head amplifies least, and imposes no constraint on layers below the final one.

  2. Noise properties: Output-space extrapolation uses a sampled advantage Aλ(v)=ℓ(v)−(λ−1)ρ(v)A_\lambda(v) = \ell(v) - (\lambda-1)\rho(v) where ℓ(v)=log⁡p(v)−log⁡q(v)\ell(v) = \log p(v) - \log q(v) and ρ(v)=log⁡q(v)−log⁡b(v)\rho(v) = \log q(v) - \log b(v).

Proposition 4.1 (Conditional variance under extrapolation):

Varv∼p[Aλ(v)]=Varv∼p[ℓ(v)]+(λ−1)2Varv∼p[ρ(v)]−2(λ−1)Covv∼p[ℓ(v),ρ(v)]\text{Var}_{v \sim p}\left[A_\lambda(v)\right] = \text{Var}_{v \sim p}[\ell(v)] + (\lambda-1)^2 \text{Var}_{v \sim p}[\rho(v)] - 2(\lambda-1)\text{Cov}_{v \sim p}[\ell(v), \rho(v)]

The second term does not vanish as p→qp \to q unless log⁡q−log⁡b\log q - \log b is constant on the support of qq. In contrast, the RIDE gradient is deterministic given a rollout, so its conditional variance is zero.

Optimization Interpretation

Proposition 4.2 (Reward-and-penalty form): Minimizing Lλ\mathcal{L}_\lambda is equivalent to maximizing:

(λ−1)r(h)−12∥h−hT∥22(\lambda - 1) r(h) - \frac{1}{2}\|h - h_T\|_2^2

where r(h)=⟨h−hT,Δ⟩r(h) = \langle h - h_T, \Delta \rangle. The reward rr is the inner product of the student's displacement from the teacher with the RL-induced residual, and the penalty is quadratic in the same displacement. The maximizer is h⋆=hT+(λ−1)Δh^{\star} = h_T + (\lambda-1)\Delta.


Empirical Validation / Results

Experimental Setup

  • Models: Four base/RL-teacher pairs: R1-Distill-1.5B → JustRL-1.5B, Qwen3-4B → Just-Qwen3-4B, Llama-3.2-3B → Just-Llama-3.2-3B, Phi-4-mini → Just-Phi-4-mini.
  • Training: Prompts from DAPO-Math-17K, 500 training steps, 3 independent seeds (14, 42, 2027).
  • Evaluation: Avg@16 on AIME24, AIME25, and AIMO (AMC 2022–2023).
  • Baselines: OPD (top-1 and top-16), OPRD (representation matching, λ=1\lambda=1), ExOPD (output-space extrapolation, λ=1.25\lambda=1.25).

Main Results (Table 1)

MethodR1-Distill-1.5BQwen3-4BLlama-3.2-3BPhi-4-mini
Teacher55.3065.5913.0117.96
Student (untouched)39.0062.826.9915.13
OPD top-152.9059.241.0812.90
OPD top-1652.5359.745.9716.76
OPRD54.5062.0110.9417.33
ExOPD49.8748.756.9512.22
RIDE56.3866.0713.3318.30

Key findings:

  • RIDE is the only method whose mean lies above the teacher line on all four pairs.
  • RIDE outperforms OPRD by 0.97–4.06 points, isolating the benefit of residual extrapolation.
  • ExOPD falls below its teacher on every pair and below the untouched student on three pairs (by 14.1 points on Qwen3-4B), confirming the noise amplification predicted by Proposition 4.1.

Coefficient Sweep (Figure 4)

  • RIDE: Final Avg@16 rises with λ\lambda from 48.2 at λ=0.5\lambda=0.5 to 54.3 at λ=1\lambda=1, peaking at 55.4 at λ=1.25\lambda=1.25. Every λ∈[1.15,1.35]\lambda \in [1.15, 1.35] improves over OPRD. Degradation beyond λ=1.35\lambda=1.35 is graceful (λ=2\lambda=2 finishes at 52.4).
  • ExOPD: Best at λ≤1\lambda \leq 1 (52.9), harmed by every λ>1\lambda > 1. The λ=1.25\lambda=1.25 run peaks early then declines to 49.9; λ=2\lambda=2 collapses to 46.2 with format score falling from 92% to 64%.

Mechanistic Analysis (Figure 5)

  • Head attenuation: The residual retains only 0.59 of the head gain of an equal-norm isotropic direction. The 512 weakest head directions hold 79.8% of the residual's hidden-state energy (vs. 33.3% under isotropy). These directions receive 73.0% of the representation-loss gradient but only 48.6% of the output-KL gradient.
  • Student alignment: The student's update aligns with the residual (cosine 0.954), with a projection onto the residual of 1.69, above the target coefficient of 1.25 — the student continues past the target along the same direction.
  • Conditional variance: The ExOPD conditional variance rises 12.6× from λ=1\lambda=1 to λ=2\lambda=2, with the extrapolation term dominating beyond λ≈1.3\lambda \approx 1.3. The RIDE per-position gradient norm is invariant to next-token resampling (variance = 0).

Direction Ablation (Table 3)

Holding displacement magnitude fixed and altering only direction:

  • Random direction: 55.04 (vs. OPRD 54.50)
  • Reversed direction (λ=0.75\lambda=0.75): 53.90
  • Mismatched origin (different base model): 54.75
  • Trajectory-mismatched residual: 55.12
  • RIDE (RL-induced residual): 56.38

The RL-induced residual itself is responsible for RIDE's gain.


Theoretical and Practical Implications

Theoretical Significance

  1. Representation-space extrapolation as reward maximization: RIDE formalizes the notion that RL-induced representation residuals encode a usable "direction" for improvement, providing a representation-space counterpart to output-space reward extrapolation with a deterministic gradient.

  2. Diagnosis of output-space extrapolation failure: Proposition 4.1 provides a rigorous explanation for why output-space extrapolation destabilizes training: the sampled-advantage noise scales with (λ−1)2(\lambda-1)^2 regardless of signal strength, and the damage is largest when the teacher-to-base gap is small.

  3. Head attenuation quantification: The paper provides concrete measurements showing how the language-model head anisotropically attenuates representation changes, with the residual concentrating in the weakest head directions.

Practical Implications

  • RIDE requires only the pre-RL checkpoint and a shared representation space, making it a practical drop-in replacement for OPRD pipelines with a single additional frozen forward pass.
  • The method is stable across a range of extrapolation coefficients (λ∈[1.15,1.35]\lambda \in [1.15, 1.35]), unlike output-space extrapolation which degrades sharply.
  • The transfer volume to the trainer equals that of OPRD (a single hidden-state tensor), with no additional communication overhead.

Conclusion

Main Takeaways

  • An RL-trained teacher supplies a direction as well as a destination; the layerwise hidden-state difference between teacher and pre-RL checkpoint on identical prefixes measures what RL changed.
  • RIDE displaces the representation-matching target beyond the teacher along this residual with a single coefficient λ\lambda, recovering OPRD exactly at λ=1\lambda=1.
  • Measuring the displacement before the head matters because the head attenuates the residual anisotropically and leaves earlier layers unconstrained.
  • RIDE approaches or exceeds its RL-trained teacher on all four base/teacher pairs, consistently outperforms OPRD and output-space extrapolation, and remains stable where the latter collapses.

Limitations and Future Work

  • RIDE requires the pre-RL checkpoint and a shared representation space.
  • Uses a single global λ\lambda (no per-layer or per-position adaptation).
  • Evaluated only on mathematical reasoning with one RL recipe.
  • Future work: relaxing these constraints, studying the safety and calibration of students trained on hidden-state targets, and exploring per-layer or adaptive extrapolation coefficients.

Related papers