# Scaling Properties of Same-Family On-Policy Distillation

> On-policy distillation transfers RL expertise across model scales predictably, with weak-to-strong students exceeding teacher performance and scaling laws forecasting peak gold scores within one accuracy point.

- **Source:** [arXiv](https://arxiv.org/abs/2609.32722)
- **Published:** 2026-10-01
- **Permalink:** https://picx.dev/p/YcKnf7
- **Whiteboard:** https://picx.dev/p/YcKnf7/image

## Summary

## Summary

## Summary (Overview)
- This paper systematically studies the **scaling properties of on-policy distillation (OPD)** for transferring reinforcement learning (RL) expertise across model scales in math reasoning tasks.
- The authors identify a **regular "useful-transfer" regime** where gold score (G) rises approximately linearly with $d = \sqrt{KL(\pi_\theta \| \pi_{ref})}$, the square root of token-level reverse KL divergence from the student initialization.
- **Weak-to-strong transfer works remarkably well**: compact RL experts can transfer capability to much larger students, with peak gold scores exceeding the teacher's own in every observed weak-to-strong pair.
- The paper fits **joint power laws** in student scale, teacher scale, and teacher gold score that predict peak gold score ($G_{peak}$) and useful-transfer slope, extrapolating to held-out scales within one accuracy point.
- Two OPD variants are compared: **Vanilla-OPD** and **Delta-OPD**, with findings that bootstrapping weak-to-strong OPD does not improve on direct transfer, and off-policy cold starts harm weak-to-strong transfer.

## Introduction and Theoretical Foundation

The paper addresses a fundamental question in LLM post-training: **how much RL-acquired capability transfers across model scales, and how quickly?**

**Background**: Modern reasoning models acquire task expertise through reinforcement learning. On-policy distillation (OPD) transfers this expertise between models—a teacher policy provides token-level supervision on rollouts sampled from the student, avoiding the exposure bias of offline teacher data.

**Key Motivation**: The central research question is whether OPD outcomes can be estimated from teacher and student scale *before* training, analogous to how reward model overoptimization studies characterize gold reward dynamics as functions of KL divergence from initialization.

**Theoretical Foundation**: The paper formalizes three distillation objectives:
1. **Vanilla-OPD**: Minimizes sequence-level reverse KL divergence between policy and teacher:
$$\min_{\theta} \mathcal{L}_{V}(\theta) = \min_{\theta} \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_{\theta}(\cdot|x)} \left[ \text{KL}(\pi_{\theta}(y|x) \| \pi_{T}(y|x)) \right] \tag{1}$$

2. **Delta-OPD**: Derives token rewards from the policy shift the teacher acquired during RL:
$$A_{t}^{\Delta}(x, y) = r_{t}^{\Delta} := \log \pi_{T}(y_t | x, y_{<t}) - \log \pi_{T}^{\text{base}}(y_t | x, y_{<t})$$

3. **Off-policy distillation (OfPD)**: SFT on teacher demonstrations (used as a baseline/comparison).

## Methodology

**Models**: Qwen2.5 Base models (not instruction-tuned) at 0.5B, 1.5B, 3B, 7B, and 14B parameters.

**Training Setup**:
- Teachers obtained via GRPO RL on mixed GSM8K + MATH training split (14.8K examples)
- Evaluation on corresponding mixed test split (6.3K examples)
- OPD and RL use the same training prompts
- Each OPD run: at most 10 epochs (580 updates) with periodic held-out evaluation

**Experimental Design**:
- **25 teacher–student combinations** spanning weak-to-strong, same-base, and strong-to-weak setups
- **Training progress metric**: $d := \sqrt{k_3}$ where $k_3$ is token-mean KL (square root of quadratic KL measure)

**Key Comparisons**:
1. **Pure OPD** vs. **Delta-OPD** (17 matched pairs)
2. **Pure OPD** vs. **OfPD** (off-policy supervision)
3. **Direct weak-to-strong OPD** vs. **bootstrapped** transfer through intermediate scales

## Empirical Validation / Results

### 1. Two-Phase Training Dynamics
OPD dynamics form a **two-phase process**:
- **Useful-transfer regime**: Gold score rises approximately linearly with $d$ at rate $m$
- **Noisy tail**: Attenuated improvement, saturation, or regression (unlike the regular overoptimization regime of PPO)

### 2. Scaling Laws

**Peak Gold Score** ($G_{peak}$):
- Joint power laws in student parameters, effective teacher size, and measured teacher score
- Validation: conditioning on teacher score **halves the leave-one-scale-out RMSE** compared to scale-only baselines
- Extrapolates to held-out largest student/teacher within **0.7 (Vanilla-OPD) and 0.4 (Delta-OPD) accuracy points**

**Table 2: Validation of peak laws** (accuracy points)

| Method | Scale-only RMSE | Joint law RMSE | Extrapolation error (largest held-out) |
|--------|----------------|----------------|----------------------------------------|
| Vanilla-OPD | (baseline) | **0.5×** | **0.7** |
| Delta-OPD | (baseline) | **0.5×** | **0.4** |

### 3. Weak-to-Strong Transfer
- **Every observed weak-to-strong pair**: student's peak gold score exceeds teacher's own
- Margins shrink as teacher scale approaches student scale
- Example: 0.5B expert teaching 14B student shows substantial gains

### 4. Variant Comparisons
- **Vanilla-OPD same-base** peaks within 0.5 accuracy points of direct-RL reference at all five scales (exceeding at three)
- **OfPD underperforms OPD** in every cell: by **28.4 points** in extreme weak-to-strong, <1 point in same-base cells

### 5. Off-Policy Cold Start Effects
- One epoch of SFT on teacher rollouts is increasingly harmful with weak-to-strong gap:
  - 3B student: **6.6 points** loss
  - 7B student: **15.7 points** loss  
  - 14B student: **19.4 points** loss
- Root cause: cold start pins students near teacher's score, erasing 32 points of 14B student's initial capability

### 6. Bootstrapping Results
- Bootstrapped weak-to-strong OPD does **not** improve on direct transfer from the smallest expert
- Direct RL on the student remains above every weak-teacher variant

## Theoretical and Practical Implications

**Theoretical Contributions**:
1. **Predictability**: OPD outcomes can be estimated from student scale, teacher scale, and teacher score *before* training—enabling principled resource allocation
2. **Multiplicative error model**: Scaling the student removes a constant fraction of whatever error teacher-induced supervision leaves, encoding scale-teacher interactions
3. **KL budget interpretation**: The transfer extent ($d_{transfer}$) reads as an approximately scale-free KL budget rather than a scaling target

**Practical Implications**:
1. **Amortized expertise**: RL expertise can be trained once at small scale and transferred predictably across a model family
2. **Design guidance**: 
   - Pure on-policy supervision is critical (OfPD severely underperforms)
   - Avoid off-policy SFT cold starts for weak-to-strong transfer
   - Direct transfer beats bootstrapping through intermediate scales
3. **Cost savings**: Enables choosing optimal teacher-student pairs without exhaustive experimentation

## Conclusion

This work establishes that **OPD scaling is predictable and regular** in its initial phase, with gold score rising linearly in $d = \sqrt{KL}$ until a noisy tail. The fitted joint power laws in student scale, teacher scale, and teacher gold score enable **accurate prediction of peak performance** before training, with extrapolation errors under one accuracy point for the largest held-out configurations.

**Key findings**:
- Weak-to-strong transfer via OPD consistently exceeds teacher performance
- Delta-OPD and Vanilla-OPD show similar scaling behavior
- On-policy supervision is essential; off-policy cold starts are increasingly harmful with capability gaps
- Bootstrapping adds no value over direct transfer

**Future Directions**:
- Extending analysis beyond math reasoning to other domains
- Investigating whether the useful-transfer regime can be extended beyond its current duration
- Exploring adaptive KL budgets that maintain the linear regime longer
- Applying these scaling laws to other distillation objectives and architectures

---

_Markdown view of https://picx.dev/p/YcKnf7, served by PicX — AI-generated visual whiteboard summaries of research papers._
