# The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

> RIDE extrapolates RL-induced hidden-state residuals beyond the teacher, consistently surpassing it across four model pairs where output-space extrapolation fails.

- **Source:** [arXiv](https://arxiv.org/abs/2609.36484)
- **Published:** 2026-10-02
- **Permalink:** https://picx.dev/p/zw3X3D
- **Whiteboard:** https://picx.dev/p/zw3X3D/image

## Summary

## Summary (Overview)

- **Problem**: On-policy distillation (OPD) aims to transfer capabilities from a reinforcement learning (RL)-trained teacher to a student model sharing the same initialization. While generalized variants allow students to *surpass* the teacher via output-space extrapolation, this approach suffers from anisotropic attenuation of representation changes by the language-model head and noise amplification from sampled-token log-probability ratios.
- **Key Insight**: RL shifts a model's internal representations relative to its base checkpoint, and this shift's *direction* can be measured at every layer. This direction, not just the destination, can be extrapolated.
- **Proposed Method**: **RIDE (RL-Induced Direction Extrapolation)** — extrapolates the RL-induced change directly in *representation space*. At every layer and token position, RIDE computes the residual between the teacher and its pre-RL checkpoint and regresses the student's hidden states toward targets displaced *beyond* the teacher along this residual.
- **Main Results**: Across four base/RL-teacher pairs (spanning different scales, architectures, and pre-training lineages), RIDE approaches or exceeds the RL-trained teacher on every pair, consistently outperforming output-space extrapolation (ExOPD) and representation-matching (OPRD) baselines.
- **Theoretical Contribution**: RIDE's objective is equivalent to maximizing a linear directional reward under a quadratic penalty centered at the teacher, and its gradient is deterministic given a rollout, unlike the noisy sampled-token advantage of output-space extrapolation.

---

## Introduction and Theoretical Foundation

### Background and Motivation

- **On-policy distillation (OPD)** trains a student to match the teacher's next-token distributions on the student's own trajectories, reducing exposure bias and providing dense learning signals. It has become a standard post-training stage for language models.
- A common use case is transferring the gains of an RL-trained teacher back to a student that shares its initialization (e.g., merging domain experts or distilling an expensive RL run).
- **Beyond-the-teacher learning**: Standard OPD treats the teacher as the endpoint. However, the teacher-to-reference log-probability ratio acts as an implicit reward; up-weighting it can yield students that outperform the teacher. Because RL updates are small, structured, and consistent across runs, an RL-trained teacher supplies a *direction* as well as a destination.

### The Problem with Output-Space Extrapolation

- **Head attenuation**: The language-model head reweights representation changes according to its anisotropic singular spectrum. The 512 weakest directions of the head carry 79.8% of the teacher-to-base residual's hidden-state energy but only 30.0% of its centered-logit energy. Thus, output-space objectives supervise much of the residual only weakly and leave earlier layers unconstrained.
- **Noise amplification**: The sampled-token log-ratio is length-biased and prone to overoptimization. Scaling it by a global coefficient amplifies extreme rewards and destabilizes training.

### Theoretical Foundation

- **Same-initialization setting**: The teacher is obtained by RL from a base checkpoint ($\pi_T = \text{RL}(\pi_B)$), and the student is initialized at that checkpoint ($\pi_{\theta_0} = \pi_B$). All three models share architecture, tokenizer, and representational origin.
- **RL-induced residual**: On a common prefix, the difference $s_{T,t} - s_{B,t}$ measures what RL changed on that context — in output space it is the implicit reward $\log \pi_T - \log \pi_B$; in representation space it is a vector at every layer, taken *before* the head.

---

## Methodology

### The RIDE Objective

The core idea is to displace the training target from the teacher along the RL-induced change:

$$\tau_t(\lambda) = s_{T,t} + (\lambda - 1)\left(s_{T,t} - s_{B,t}\right) = \lambda s_{T,t} + (1 - \lambda) s_{B,t}, \quad \lambda \geq 0$$

- At $\lambda = 1$, the target is the teacher (recovering standard OPD/OPRD).
- For $\lambda > 1$, the target lies *beyond* the teacher, continuing the change RL initiated.

**Applying this to hidden states** yields the RIDE target:

$$h^{(l)\star}_{t} = h^{(l)}_{T,t} + (\lambda - 1)\Delta^{(l)}_{t} = \lambda h^{(l)}_{T,t} + (1 - \lambda) h^{(l)}_{B,t}$$

where $\Delta^{(l)}_{t} = h^{(l)}_{T,t} - h^{(l)}_{B,t}$ is the RL-induced residual at layer $l$ and position $t$.

**Training objective** (representation-space regression):

$$\mathcal{L}_{\lambda}(\theta) = \mathbb{E}_{x, \hat{y} \sim \pi_{\theta}}\left[\frac{1}{|\mathcal{L}_{\text{layer}}|} \sum_{l \in \mathcal{L}_{\text{layer}}} \frac{1}{M} \sum_{t=1}^{T} m_t \frac{1}{d} \left\| h^{(l)}_{\theta,t} - \text{sg}\left(h^{(l)\star}_{t}\right) \right\|_2^2\right]$$

where $M = \sum_t m_t$ and $\text{sg}$ stops gradients through the target. All $L$ layers and the final $k$ response positions are supervised.

### Why Hidden States?

1. **Head attenuation**: The output-space target is the head-projected image of the RIDE target: $W^{\text{head}}(\tilde{h}_T + (\lambda-1)\tilde{\Delta}) = z_T + (\lambda-1)(z_T - z_B)$. However, a logit target constrains hidden states only weakly along directions the head amplifies least, and imposes no constraint on layers below the final one.

2. **Noise properties**: Output-space extrapolation uses a sampled advantage $A_\lambda(v) = \ell(v) - (\lambda-1)\rho(v)$ where $\ell(v) = \log p(v) - \log q(v)$ and $\rho(v) = \log q(v) - \log b(v)$.

**Proposition 4.1 (Conditional variance under extrapolation):**

$$\text{Var}_{v \sim p}\left[A_\lambda(v)\right] = \text{Var}_{v \sim p}[\ell(v)] + (\lambda-1)^2 \text{Var}_{v \sim p}[\rho(v)] - 2(\lambda-1)\text{Cov}_{v \sim p}[\ell(v), \rho(v)]$$

The second term does not vanish as $p \to q$ unless $\log q - \log b$ is constant on the support of $q$. In contrast, the RIDE gradient is deterministic given a rollout, so its conditional variance is zero.

### Optimization Interpretation

**Proposition 4.2 (Reward-and-penalty form):** Minimizing $\mathcal{L}_\lambda$ is equivalent to maximizing:

$$(\lambda - 1) r(h) - \frac{1}{2}\|h - h_T\|_2^2$$

where $r(h) = \langle h - h_T, \Delta \rangle$. The reward $r$ is the inner product of the student's displacement from the teacher with the RL-induced residual, and the penalty is quadratic in the same displacement. The maximizer is $h^{\star} = h_T + (\lambda-1)\Delta$.

---

## Empirical Validation / Results

### Experimental Setup

- **Models**: Four base/RL-teacher pairs: R1-Distill-1.5B → JustRL-1.5B, Qwen3-4B → Just-Qwen3-4B, Llama-3.2-3B → Just-Llama-3.2-3B, Phi-4-mini → Just-Phi-4-mini.
- **Training**: Prompts from DAPO-Math-17K, 500 training steps, 3 independent seeds (14, 42, 2027).
- **Evaluation**: Avg@16 on AIME24, AIME25, and AIMO (AMC 2022–2023).
- **Baselines**: OPD (top-1 and top-16), OPRD (representation matching, $\lambda=1$), ExOPD (output-space extrapolation, $\lambda=1.25$).

### Main Results (Table 1)

| Method | R1-Distill-1.5B | Qwen3-4B | Llama-3.2-3B | Phi-4-mini |
|--------|----------------|----------|--------------|------------|
| Teacher | 55.30 | 65.59 | 13.01 | 17.96 |
| Student (untouched) | 39.00 | 62.82 | 6.99 | 15.13 |
| OPD top-1 | 52.90 | 59.24 | 1.08 | 12.90 |
| OPD top-16 | 52.53 | 59.74 | 5.97 | 16.76 |
| OPRD | 54.50 | 62.01 | 10.94 | 17.33 |
| ExOPD | 49.87 | 48.75 | 6.95 | 12.22 |
| **RIDE** | **56.38** | **66.07** | **13.33** | **18.30** |

**Key findings**:
- RIDE is the **only method whose mean lies above the teacher line on all four pairs**.
- RIDE outperforms OPRD by 0.97–4.06 points, isolating the benefit of residual extrapolation.
- ExOPD falls below its teacher on every pair and below the untouched student on three pairs (by 14.1 points on Qwen3-4B), confirming the noise amplification predicted by Proposition 4.1.

### Coefficient Sweep (Figure 4)

- **RIDE**: Final Avg@16 rises with $\lambda$ from 48.2 at $\lambda=0.5$ to 54.3 at $\lambda=1$, peaking at **55.4 at $\lambda=1.25$**. Every $\lambda \in [1.15, 1.35]$ improves over OPRD. Degradation beyond $\lambda=1.35$ is graceful ($\lambda=2$ finishes at 52.4).
- **ExOPD**: Best at $\lambda \leq 1$ (52.9), harmed by every $\lambda > 1$. The $\lambda=1.25$ run peaks early then declines to 49.9; $\lambda=2$ collapses to 46.2 with format score falling from 92% to 64%.

### Mechanistic Analysis (Figure 5)

- **Head attenuation**: The residual retains only 0.59 of the head gain of an equal-norm isotropic direction. The 512 weakest head directions hold 79.8% of the residual's hidden-state energy (vs. 33.3% under isotropy). These directions receive 73.0% of the representation-loss gradient but only 48.6% of the output-KL gradient.
- **Student alignment**: The student's update aligns with the residual (cosine 0.954), with a projection onto the residual of 1.69, above the target coefficient of 1.25 — the student continues past the target along the same direction.
- **Conditional variance**: The ExOPD conditional variance rises 12.6× from $\lambda=1$ to $\lambda=2$, with the extrapolation term dominating beyond $\lambda \approx 1.3$. The RIDE per-position gradient norm is invariant to next-token resampling (variance = 0).

### Direction Ablation (Table 3)

Holding displacement magnitude fixed and altering only direction:
- Random direction: 55.04 (vs. OPRD 54.50)
- Reversed direction ($\lambda=0.75$): 53.90
- Mismatched origin (different base model): 54.75
- Trajectory-mismatched residual: 55.12
- **RIDE (RL-induced residual): 56.38**

The RL-induced residual itself is responsible for RIDE's gain.

---

## Theoretical and Practical Implications

### Theoretical Significance

1. **Representation-space extrapolation as reward maximization**: RIDE formalizes the notion that RL-induced representation residuals encode a usable "direction" for improvement, providing a representation-space counterpart to output-space reward extrapolation with a *deterministic* gradient.

2. **Diagnosis of output-space extrapolation failure**: Proposition 4.1 provides a rigorous explanation for why output-space extrapolation destabilizes training: the sampled-advantage noise scales with $(\lambda-1)^2$ regardless of signal strength, and the damage is largest when the teacher-to-base gap is small.

3. **Head attenuation quantification**: The paper provides concrete measurements showing how the language-model head anisotropically attenuates representation changes, with the residual concentrating in the weakest head directions.

### Practical Implications

- RIDE requires only the pre-RL checkpoint and a shared representation space, making it a practical drop-in replacement for OPRD pipelines with a single additional frozen forward pass.
- The method is stable across a range of extrapolation coefficients ($\lambda \in [1.15, 1.35]$), unlike output-space extrapolation which degrades sharply.
- The transfer volume to the trainer equals that of OPRD (a single hidden-state tensor), with no additional communication overhead.

---

## Conclusion

### Main Takeaways

- An RL-trained teacher supplies a **direction** as well as a destination; the layerwise hidden-state difference between teacher and pre-RL checkpoint on identical prefixes measures what RL changed.
- RIDE displaces the representation-matching target beyond the teacher along this residual with a single coefficient $\lambda$, recovering OPRD exactly at $\lambda=1$.
- Measuring the displacement **before the head** matters because the head attenuates the residual anisotropically and leaves earlier layers unconstrained.
- RIDE approaches or exceeds its RL-trained teacher on all four base/teacher pairs, consistently outperforms OPRD and output-space extrapolation, and remains stable where the latter collapses.

### Limitations and Future Work

- RIDE requires the pre-RL checkpoint and a shared representation space.
- Uses a single global $\lambda$ (no per-layer or per-position adaptation).
- Evaluated only on mathematical reasoning with one RL recipe.
- Future work: relaxing these constraints, studying the safety and calibration of students trained on hidden-state targets, and exploring per-layer or adaptive extrapolation coefficients.

---

_Markdown view of https://picx.dev/p/zw3X3D, served by PicX — AI-generated visual whiteboard summaries of research papers._
