# Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

> Domain-normalized multi-teacher distillation, which rescales teacher feedback by log-ratio dispersion, outperforms label-routed MOPD by up to 3.08 points across model scales.

- **Source:** [arXiv](https://arxiv.org/abs/2609.35347)
- **Published:** 2026-09-30
- **Permalink:** https://picx.dev/p/mUSEiN
- **Whiteboard:** https://picx.dev/p/mUSEiN/image

## Summary

## Summary (Overview)

- **Problem**: Multi-teacher on-policy distillation (MOPD) routes each prompt to a domain-specialist teacher, but fails to calibrate *how strongly* each teacher's feedback influences the shared student. This leads to imbalanced gradients where instruction-following feedback dominates and mathematics skills fail to transfer.
- **Key Finding**: In Qwen3.5 models at 9B, 4B, and 2B scales, label-routed MOPD does not outperform a student trained by the best single specialist, and retains only 14–32% of the mathematics expert's gain at an 8K evaluation budget.
- **Proposed Solution**: **Domain-Normalized MOPD (DN-MOPD)** rescales each domain's distillation advantages by the ratio of pooled to per-domain log-ratio standard deviation, using a bounded multiplier $\text{clip}(\sigma_{\text{all}}/\sigma_d, 0.25, 4)$.
- **Results**: DN-MOPD improves the six-task average over MOPD by **1.17–2.36 percentage points at 16K** and **2.47–3.08 at 8K** evaluation budgets, across three random seeds and all model sizes.
- **Attribution**: Controls show the gain comes primarily from *reducing* the influence of dispersed instruction-following feedback rather than amplifying mathematics alone.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Post-training of large language models increasingly relies on **domain-specific reinforcement learning (RL)** with verifiable rewards: mathematical reasoning (verifiable rewards), code execution feedback, and constraint checking for instruction following. These pipelines differ in data, rewards, and optimization, so they are typically run separately from a shared initialization, producing **specialists** that each excel in one domain.

**Multi-teacher on-policy distillation (MOPD)** integrates these specialists in policy space: the student samples responses, each prompt is routed to its domain expert, and that expert's token-level log-probabilities on the student's own response provide a dense distillation signal. MOPD has been adopted in the post-training of several frontier models (DeepSeek-AI, NVIDIA, Xiaomi, etc.).

### The Core Problem: Feedback Imbalance

The routing rule specifies *which* expert supervises a prompt, but **not how strongly** its feedback moves the shared student. The paper's empirical investigation reveals:

- Token-level teacher–student log-ratios from the instruction-following expert are **2.3 to 4.4 times as dispersed** as the pooled batch signal
- Mathematics feedback is about **half as dispersed** as the pooled signal
- For the initial 4B student, the instruction-following loss supplies **94% of the combined gradient** under equal weights

> "Routing determines where feedback comes from, but not how much it counts."

### Theoretical Foundation: Distillation Advantages

The token-level distillation advantage is defined as:

$$A_t = \log p_{T_d}(y_t \mid h_t) - \log \pi_u(y_t \mid h_t)$$

where $T_d$ is the domain specialist teacher, $\pi_u$ is the student, and $h_t = (x, y_{<t})$ is the preceding context. A positive $A_t$ encourages the sampled token; a negative $A_t$ discourages it. The **magnitude** weights that token's policy-gradient contribution.

---

## Methodology

### Domain Normalization Rule

For each batch, DN-MOPD measures the spread of teacher–student log-ratios within each domain and rescales:

1. Let $r_{i,t} = \ell_{i,t}^T - \ell_{i,t}^{\text{roll}}$ be the token-level log-ratio (teacher minus student rollout log-probability)
2. Compute $\sigma_d$ = population standard deviation of $r_{i,t}$ over valid response tokens with domain $d$
3. Compute $\sigma_{\text{all}}$ over all valid response tokens pooled across domains
4. Apply the multiplier:

$$w_d = \text{clip}\left(\frac{\sigma_{\text{all}}}{\sigma_d}, 0.25, 4\right), \quad \tilde{A}_t = w_d A_t$$

For a non-degenerate domain whose ratio is not clipped, the scaled rollout signal satisfies:

$$\text{Std}(w_d r_d) = \sigma_{\text{all}}$$

This **amplifies feedback with smaller spread** (e.g., mathematics) and **attenuates feedback with larger spread** (e.g., instruction following).

### Student Update

The actor recomputes pre-update student log-probabilities $\ell^{\text{actor}}$ on sampled responses, forming $A_{i,t} = \ell_{i,t}^T - \ell_{i,t}^{\text{actor}}$. The scaled advantage is detached and used in the clipped OPD loss:

$$\mathcal{L}_{i,t}(\theta) = -\min\left\{\rho_{i,t}(\theta)\tilde{A}_{i,t}, \text{clip}\left(\rho_{i,t}(\theta), 1-\eta_-, 1+\eta_+\right)\tilde{A}_{i,t}\right\}$$

where $\rho_{i,t}(\theta) = \pi_\theta(y_{i,t} \mid h_{i,t}) / \exp(\ell_{i,t}^{\text{actor}})$ is the importance ratio, and $\eta_-$, $\eta_+$ are policy-ratio clipping bounds.

### Algorithm 1: DN-MOPD, one training iteration

```
Require: student π_u, teachers {T_d}, labeled prompts
1: Sample responses y_i ~ π_u(· | x_i)
2: Score each y_i with its domain teacher T_{d_i}
3: r_{i,t} ← ℓ_{i,t}^T - ℓ_{i,t}^roll on valid tokens
4: σ_all ← Std({r_{i,t}})
5: for each domain d do
6:   σ_d ← Std({r_{i,t} : d_i = d})
7:   w_d ← clip(σ_all / σ_d, 0.25, 4)
8:   Use w_d = 1 if either statistic is degenerate
9: end for
10: Recompute A_{i,t} ← ℓ_{i,t}^T - ℓ_{i,t}^actor
11: Ã_{i,t} ← stopgrad(w_{d_i} A_{i,t})
12: Update π_u on Ã with the clipped OPD objective
```

### Key Design Choices

- **Sign preservation**: Multiplies by a positive factor without subtracting the domain mean, preserving whether each advantage encourages or discourages the sampled token
- **Bounded adjustment**: The [0.25, 4] bounds limit extreme adjustments (the IF multiplier typically sits at the 0.25 floor)
- **Zero additional cost**: Reuses log-ratios already available from rollout scoring—no additional teacher call, teacher model, or learned router
- **MOPD as special case**: Setting every $w_d = 1$ recovers standard label-routed MOPD

---

## Empirical Validation / Results

### Experimental Setup

- **Models**: Qwen3.5 at 9B, 4B, and 2B with independently trained specialist pools (mathematics, code, instruction following)
- **Evaluation**: Six public benchmarks—AIME25/AIME26 (math), LiveCodeBench v5/v6 (code), IFEval/IFBench (instruction following); sampling 64/6/16 responses per question respectively
- **Training**: 80 updates, student seed 42 (seeds 43, 44 for robustness), 8,192-token training response cap; evaluation at both 8K and 16K token caps

### Main Results

**Table 1: Capability integration at 9B** (selected rows)

| Method | Math (avg@64) | Code (avg@6) | IF (avg@16) | Total |
|--------|--------------|-------------|-------------|-------|
| Initial student | 60.2 | 53.2 | 58.1 | 57.1 |
| Label-routed MOPD | 59.3 | 54.7 | 61.2 | 58.4 |
| **DN-MOPD (Ours)** | **63.3** | **53.9** | **61.6** | **59.6** |

**Table 2: Capability integration at smaller scales** (Total scores)

| Method | 4B Total | 2B Total |
|--------|----------|----------|
| Initial student | 47.7 | 24.0 |
| Label-routed MOPD | 50.3 | 26.6 |
| **DN-MOPD (Ours)** | **52.5** | **29.0** |

### Key Findings

1. **DN-MOPD consistently beats Label-routed MOPD** at every size and both evaluation budgets, with paired confidence intervals above zero
2. **The main gain is in mathematics**: At 16K, Label shows no math gain over the initial student at any size, while DN-MOPD improves math at every size
3. **DN-MOPD exceeds the strongest single-teacher student** at every size (though some intervals include zero)
4. **Gains are not from longer generations**: DN-MOPD produces *shorter* math answers with fewer responses hitting the generation cap (Table 4)

### Ablation and Control Experiments

**Table 3: Changes to label-routed OPD** (Total point differences vs. Label at 16K)

| Variant | Teacher rule | Signal change | 9B | 4B | 2B |
|---------|-------------|---------------|-----|-----|-----|
| Uniform pool | Mixture | Unchanged | −0.41 | −0.69 | +0.30 |
| Dynamic router | Response | Unchanged | −0.25 | +0.28 | −0.07 |
| Annealed injection | Domain | Early imitation | −0.24 | +1.11 | +0.30 |
| Math ×2 only | Domain | Math scale | — | +1.20 | +0.27 |
| IF ×0.25 only | Domain | IF scale | — | +1.79 | +2.19 |
| **DN-MOPD** | Domain | Domain scale | **+1.17** | **+2.24** | **+2.36** |

**Key attribution**: Reducing IF feedback alone recovers most of DN-MOPD's improvement (+1.79 at 4B, +2.19 at 2B), while amplifying math alone recovers only about half at 4B and little at 2B.

### Training Duration Analysis

**Table 5: Training-duration comparison** (Total at 16K)

| Method | 80 updates | 160 updates | Δ Total |
|--------|-----------|-------------|---------|
| **9B**: Label / DN-MOPD | 58.4 / 59.6 | 59.4 / 60.7 | +1.0 / +1.2 |
| **4B**: Label / DN-MOPD | 50.3 / 52.5 | 51.6 / 53.1 | +1.3 / +0.5 |
| **2B**: Label / DN-MOPD | 26.6 / 29.0 | 28.3 / 30.0 | +1.7 / +1.1 |

DN-MOPD remains ahead at every size when training continues to 160 updates, though the margin narrows at smaller scales.

### Feedback Scale Diagnostics

- **IF log-ratios** are 2.3–4.4 times more dispersed than the pooled signal; **math log-ratios** about half as dispersed
- DN-MOPD amplifies mathematics by **1.5–2.1×** at every size and keeps code near 1 at 4B and 2B
- The IF multiplier sits at the **0.25 floor** in most batches, so clipping *bounds* rather than equalizes that domain's scale
- IF supplies about **1% of response tokens but up to half of the pooled variance**

---

## Theoretical and Practical Implications

### Theoretical Implications

1. **Teacher assignment is necessary but not sufficient** for multi-teacher distillation. The *scale* of each teacher's feedback is an independent design choice that significantly affects integration quality.

2. **Feedback dispersion, not just mean magnitude, shapes the shared gradient**. Existing methods (e.g., Open-MOPD) allocate budget by mean absolute rewards or remaining capability gaps; DN-MOPD shows that dispersion-based calibration is a complementary and effective signal.

3. **The gradient dominance of a single domain can be extreme**: IF contributes 94% of the combined gradient for the initial 4B student under equal weights, despite representing only ~1% of tokens. This suggests that standard loss averaging can severely misallocate learning signal.

### Practical Implications

1. **Zero-cost improvement**: DN-MOPD adds one operation to MOPD with no additional teacher calls, teacher models, or learned routers—making it a drop-in replacement for existing MOPD pipelines.

2. **Fixed weights vs. adaptive estimation**: Fixed weights near DN-MOPD's measured multipliers perform comparably at 9B and 4B, but per-batch estimation adds +1.10 points at 2B. This suggests practitioners can start with fixed weights and add adaptive estimation for smaller models.

3. **Evaluation against single-teacher baselines**: The paper shows MOPD students should be evaluated against the strongest single-teacher student, not just the initial model—a stronger reference that multi-teacher methods must beat.

---

## Conclusion

### Main Takeaways

- **MOPD with label routing alone fails to transfer specialist skills effectively**: it does not beat the strongest single-teacher student and loses most of the mathematics gain.
- **DN-MOPD adds the missing control**: rescaling each domain's distillation advantages by the batch-wise spread of teacher–student log-ratios, with bounded multipliers.
- **The improvement is consistent and robust**: across three model sizes, three student seeds, two evaluation budgets, and extended training durations.
- **The mechanism is understood**: DN-MOPD primarily works by limiting the influence of dispersed instruction-following feedback, which otherwise dominates the shared gradient.

### Limitations and Future Directions

- **Single model family**: Comparisons use one expert pool per size within Qwen3.5; results may not generalize to other architectures or teacher configurations
- **Fixed clipping bounds**: The [0.25, 4] bounds were not tuned per size; the IF multiplier usually sits at the lower bound, so the rule *bounds* rather than equalizes that domain's scale
- **Seed coverage**: Three student seeds do not cover retraining the experts themselves
- **Convergence behavior**: The eventual ordering at convergence remains unresolved, though DN-MOPD remains ahead through 160 updates

### Future Work

- Investigating whether adaptive clipping bounds or domain-specific normalization schemes yield further gains
- Applying DN-MOPD to other teacher pools, model families, and domain configurations
- Exploring the interaction between feedback scale calibration and other design choices (token allocation, reward refresh, etc.)

---

_Markdown view of https://picx.dev/p/mUSEiN, served by PicX — AI-generated visual whiteboard summaries of research papers._
