Summary (Overview)

  • Problem: Multi-teacher on-policy distillation (MOPD) routes each prompt to a domain-specialist teacher, but fails to calibrate how strongly each teacher's feedback influences the shared student. This leads to imbalanced gradients where instruction-following feedback dominates and mathematics skills fail to transfer.
  • Key Finding: In Qwen3.5 models at 9B, 4B, and 2B scales, label-routed MOPD does not outperform a student trained by the best single specialist, and retains only 14–32% of the mathematics expert's gain at an 8K evaluation budget.
  • Proposed Solution: Domain-Normalized MOPD (DN-MOPD) rescales each domain's distillation advantages by the ratio of pooled to per-domain log-ratio standard deviation, using a bounded multiplier clip(σall/σd,0.25,4)\text{clip}(\sigma_{\text{all}}/\sigma_d, 0.25, 4).
  • Results: DN-MOPD improves the six-task average over MOPD by 1.17–2.36 percentage points at 16K and 2.47–3.08 at 8K evaluation budgets, across three random seeds and all model sizes.
  • Attribution: Controls show the gain comes primarily from reducing the influence of dispersed instruction-following feedback rather than amplifying mathematics alone.

Introduction and Theoretical Foundation

Background and Motivation

Post-training of large language models increasingly relies on domain-specific reinforcement learning (RL) with verifiable rewards: mathematical reasoning (verifiable rewards), code execution feedback, and constraint checking for instruction following. These pipelines differ in data, rewards, and optimization, so they are typically run separately from a shared initialization, producing specialists that each excel in one domain.

Multi-teacher on-policy distillation (MOPD) integrates these specialists in policy space: the student samples responses, each prompt is routed to its domain expert, and that expert's token-level log-probabilities on the student's own response provide a dense distillation signal. MOPD has been adopted in the post-training of several frontier models (DeepSeek-AI, NVIDIA, Xiaomi, etc.).

The Core Problem: Feedback Imbalance

The routing rule specifies which expert supervises a prompt, but not how strongly its feedback moves the shared student. The paper's empirical investigation reveals:

  • Token-level teacher–student log-ratios from the instruction-following expert are 2.3 to 4.4 times as dispersed as the pooled batch signal
  • Mathematics feedback is about half as dispersed as the pooled signal
  • For the initial 4B student, the instruction-following loss supplies 94% of the combined gradient under equal weights

"Routing determines where feedback comes from, but not how much it counts."

Theoretical Foundation: Distillation Advantages

The token-level distillation advantage is defined as:

At=log⁡pTd(yt∣ht)−log⁡πu(yt∣ht)A_t = \log p_{T_d}(y_t \mid h_t) - \log \pi_u(y_t \mid h_t)

where TdT_d is the domain specialist teacher, πu\pi_u is the student, and ht=(x,y<t)h_t = (x, y_{<t}) is the preceding context. A positive AtA_t encourages the sampled token; a negative AtA_t discourages it. The magnitude weights that token's policy-gradient contribution.


Methodology

Domain Normalization Rule

For each batch, DN-MOPD measures the spread of teacher–student log-ratios within each domain and rescales:

  1. Let ri,t=ℓi,tT−ℓi,trollr_{i,t} = \ell_{i,t}^T - \ell_{i,t}^{\text{roll}} be the token-level log-ratio (teacher minus student rollout log-probability)
  2. Compute σd\sigma_d = population standard deviation of ri,tr_{i,t} over valid response tokens with domain dd
  3. Compute σall\sigma_{\text{all}} over all valid response tokens pooled across domains
  4. Apply the multiplier:
wd=clip(σallσd,0.25,4),A~t=wdAtw_d = \text{clip}\left(\frac{\sigma_{\text{all}}}{\sigma_d}, 0.25, 4\right), \quad \tilde{A}_t = w_d A_t

For a non-degenerate domain whose ratio is not clipped, the scaled rollout signal satisfies:

Std(wdrd)=σall\text{Std}(w_d r_d) = \sigma_{\text{all}}

This amplifies feedback with smaller spread (e.g., mathematics) and attenuates feedback with larger spread (e.g., instruction following).

Student Update

The actor recomputes pre-update student log-probabilities ℓactor\ell^{\text{actor}} on sampled responses, forming Ai,t=ℓi,tT−ℓi,tactorA_{i,t} = \ell_{i,t}^T - \ell_{i,t}^{\text{actor}}. The scaled advantage is detached and used in the clipped OPD loss:

Li,t(θ)=−min⁡{ρi,t(θ)A~i,t,clip(ρi,t(θ),1−η−,1+η+)A~i,t}\mathcal{L}_{i,t}(\theta) = -\min\left\{\rho_{i,t}(\theta)\tilde{A}_{i,t}, \text{clip}\left(\rho_{i,t}(\theta), 1-\eta_-, 1+\eta_+\right)\tilde{A}_{i,t}\right\}

where ρi,t(θ)=πθ(yi,t∣hi,t)/exp⁡(ℓi,tactor)\rho_{i,t}(\theta) = \pi_\theta(y_{i,t} \mid h_{i,t}) / \exp(\ell_{i,t}^{\text{actor}}) is the importance ratio, and η−\eta_-, η+\eta_+ are policy-ratio clipping bounds.

Algorithm 1: DN-MOPD, one training iteration

Require: student π_u, teachers {T_d}, labeled prompts
1: Sample responses y_i ~ π_u(· | x_i)
2: Score each y_i with its domain teacher T_{d_i}
3: r_{i,t} ← ℓ_{i,t}^T - ℓ_{i,t}^roll on valid tokens
4: σ_all ← Std({r_{i,t}})
5: for each domain d do
6:   σ_d ← Std({r_{i,t} : d_i = d})
7:   w_d ← clip(σ_all / σ_d, 0.25, 4)
8:   Use w_d = 1 if either statistic is degenerate
9: end for
10: Recompute A_{i,t} ← ℓ_{i,t}^T - ℓ_{i,t}^actor
11: Ã_{i,t} ← stopgrad(w_{d_i} A_{i,t})
12: Update π_u on à with the clipped OPD objective

Key Design Choices

  • Sign preservation: Multiplies by a positive factor without subtracting the domain mean, preserving whether each advantage encourages or discourages the sampled token
  • Bounded adjustment: The [0.25, 4] bounds limit extreme adjustments (the IF multiplier typically sits at the 0.25 floor)
  • Zero additional cost: Reuses log-ratios already available from rollout scoring—no additional teacher call, teacher model, or learned router
  • MOPD as special case: Setting every wd=1w_d = 1 recovers standard label-routed MOPD

Empirical Validation / Results

Experimental Setup

  • Models: Qwen3.5 at 9B, 4B, and 2B with independently trained specialist pools (mathematics, code, instruction following)
  • Evaluation: Six public benchmarks—AIME25/AIME26 (math), LiveCodeBench v5/v6 (code), IFEval/IFBench (instruction following); sampling 64/6/16 responses per question respectively
  • Training: 80 updates, student seed 42 (seeds 43, 44 for robustness), 8,192-token training response cap; evaluation at both 8K and 16K token caps

Main Results

Table 1: Capability integration at 9B (selected rows)

MethodMath (avg@64)Code (avg@6)IF (avg@16)Total
Initial student60.253.258.157.1
Label-routed MOPD59.354.761.258.4
DN-MOPD (Ours)63.353.961.659.6

Table 2: Capability integration at smaller scales (Total scores)

Method4B Total2B Total
Initial student47.724.0
Label-routed MOPD50.326.6
DN-MOPD (Ours)52.529.0

Key Findings

  1. DN-MOPD consistently beats Label-routed MOPD at every size and both evaluation budgets, with paired confidence intervals above zero
  2. The main gain is in mathematics: At 16K, Label shows no math gain over the initial student at any size, while DN-MOPD improves math at every size
  3. DN-MOPD exceeds the strongest single-teacher student at every size (though some intervals include zero)
  4. Gains are not from longer generations: DN-MOPD produces shorter math answers with fewer responses hitting the generation cap (Table 4)

Ablation and Control Experiments

Table 3: Changes to label-routed OPD (Total point differences vs. Label at 16K)

VariantTeacher ruleSignal change9B4B2B
Uniform poolMixtureUnchanged−0.41−0.69+0.30
Dynamic routerResponseUnchanged−0.25+0.28−0.07
Annealed injectionDomainEarly imitation−0.24+1.11+0.30
Math ×2 onlyDomainMath scale—+1.20+0.27
IF ×0.25 onlyDomainIF scale—+1.79+2.19
DN-MOPDDomainDomain scale+1.17+2.24+2.36

Key attribution: Reducing IF feedback alone recovers most of DN-MOPD's improvement (+1.79 at 4B, +2.19 at 2B), while amplifying math alone recovers only about half at 4B and little at 2B.

Training Duration Analysis

Table 5: Training-duration comparison (Total at 16K)

Method80 updates160 updatesΔ Total
9B: Label / DN-MOPD58.4 / 59.659.4 / 60.7+1.0 / +1.2
4B: Label / DN-MOPD50.3 / 52.551.6 / 53.1+1.3 / +0.5
2B: Label / DN-MOPD26.6 / 29.028.3 / 30.0+1.7 / +1.1

DN-MOPD remains ahead at every size when training continues to 160 updates, though the margin narrows at smaller scales.

Feedback Scale Diagnostics

  • IF log-ratios are 2.3–4.4 times more dispersed than the pooled signal; math log-ratios about half as dispersed
  • DN-MOPD amplifies mathematics by 1.5–2.1× at every size and keeps code near 1 at 4B and 2B
  • The IF multiplier sits at the 0.25 floor in most batches, so clipping bounds rather than equalizes that domain's scale
  • IF supplies about 1% of response tokens but up to half of the pooled variance

Theoretical and Practical Implications

Theoretical Implications

  1. Teacher assignment is necessary but not sufficient for multi-teacher distillation. The scale of each teacher's feedback is an independent design choice that significantly affects integration quality.

  2. Feedback dispersion, not just mean magnitude, shapes the shared gradient. Existing methods (e.g., Open-MOPD) allocate budget by mean absolute rewards or remaining capability gaps; DN-MOPD shows that dispersion-based calibration is a complementary and effective signal.

  3. The gradient dominance of a single domain can be extreme: IF contributes 94% of the combined gradient for the initial 4B student under equal weights, despite representing only ~1% of tokens. This suggests that standard loss averaging can severely misallocate learning signal.

Practical Implications

  1. Zero-cost improvement: DN-MOPD adds one operation to MOPD with no additional teacher calls, teacher models, or learned routers—making it a drop-in replacement for existing MOPD pipelines.

  2. Fixed weights vs. adaptive estimation: Fixed weights near DN-MOPD's measured multipliers perform comparably at 9B and 4B, but per-batch estimation adds +1.10 points at 2B. This suggests practitioners can start with fixed weights and add adaptive estimation for smaller models.

  3. Evaluation against single-teacher baselines: The paper shows MOPD students should be evaluated against the strongest single-teacher student, not just the initial model—a stronger reference that multi-teacher methods must beat.


Conclusion

Main Takeaways

  • MOPD with label routing alone fails to transfer specialist skills effectively: it does not beat the strongest single-teacher student and loses most of the mathematics gain.
  • DN-MOPD adds the missing control: rescaling each domain's distillation advantages by the batch-wise spread of teacher–student log-ratios, with bounded multipliers.
  • The improvement is consistent and robust: across three model sizes, three student seeds, two evaluation budgets, and extended training durations.
  • The mechanism is understood: DN-MOPD primarily works by limiting the influence of dispersed instruction-following feedback, which otherwise dominates the shared gradient.

Limitations and Future Directions

  • Single model family: Comparisons use one expert pool per size within Qwen3.5; results may not generalize to other architectures or teacher configurations
  • Fixed clipping bounds: The [0.25, 4] bounds were not tuned per size; the IF multiplier usually sits at the lower bound, so the rule bounds rather than equalizes that domain's scale
  • Seed coverage: Three student seeds do not cover retraining the experts themselves
  • Convergence behavior: The eventual ordering at convergence remains unresolved, though DN-MOPD remains ahead through 160 updates

Future Work

  • Investigating whether adaptive clipping bounds or domain-specific normalization schemes yield further gains
  • Applying DN-MOPD to other teacher pools, model families, and domain configurations
  • Exploring the interaction between feedback scale calibration and other design choices (token allocation, reward refresh, etc.)

Related papers