Summary (Overview)
- Core Question: The paper investigates whether expanding alignment coverage in Cross-Tokenizer On-Policy Distillation (OPD) improves student learning, or whether supervision reliability matters more than coverage.
- Key Finding 1: Strict 1:1 token alignment already covers 85.57–96.98% of student-generated tokens across three heterogeneous teacher–student pairs, despite static vocabulary Jaccard overlap of only 39.49–64.87%.
- Key Finding 2: Adding span MSE supervision on mismatch groups achieves complete structural coverage but reduces accuracy across all 18 tested positive-weight settings (0.27–1.20 percentage points below strict baselines).
- Key Finding 3: A compact student-selected top-16 subset of the shared vocabulary retains at least 96% of the full-average improvement of full shared-vocabulary OPD, outperforming all evaluated cross-tokenizer baselines (ULD, Extended ULD, GOLD, SimCT).
- Key Finding 4: Gradient diagnostics reveal that span gradients have weak or negative directional agreement with strict gradients and grow in relative magnitude during training, explaining the accuracy drop from span supervision.
Introduction and Theoretical Foundation
Background
On-Policy Distillation (OPD) trains a student language model on its own generations using teacher feedback. The student samples trajectories from its current policy, and the teacher provides supervision at the exact prefixes the student visits. This connects teacher feedback to the student's evolving behavior.
The Cross-Tokenizer Challenge
Extending OPD across model families introduces a fundamental alignment problem: different tokenizers assign different token boundaries and use different vocabularies. The same response can be tokenized differently, and next-token distributions are defined over different vocabulary spaces. Cross-Tokenizer OPD must therefore handle alignment at two levels:
- Sequence level: Matching token boundaries between teacher and student tokenizations
- Vocabulary level: Comparing predictions across different vocabulary spaces
Key Concepts
Alignment coverage describes how much of the response and vocabularies can be compared across tokenizers. Supervision reliability concerns whether the resulting teacher targets provide useful guidance for student learning.
The paper challenges the assumption that maximizing alignment coverage necessarily improves distillation quality. Prior work (SimCT, Byte-Prefix Marginalization) has focused on recovering excluded supervision signals, but the authors question whether these recovered signals actually contribute to learning.
Theoretical Foundation
With a shared tokenizer, OPD minimizes:
This can be viewed as dense KL-constrained reinforcement learning, where the teacher distribution induces a token-level reward.
Methodology
Token-Group Alignment
For scoring and alignment, the decoded student response is tokenized with both models' tokenizers. The response is partitioned into R pairs of aligned token groups:
Groups are divided into strict 1:1 groups and mismatch groups:
Shared Vocabulary Restriction
Both distributions are restricted and renormalized over the shared vocabulary :
Strict Cross-Tokenizer Objective
Span Supervision (MSE on Mismatch Groups)
For mismatch groups, the probabilities of observed token paths are computed as:
The span loss uses MSE on log-probabilities:
Total loss combines both terms:
Experimental Setup
- Model pairs: Qwen2.5-7B-Instruct→Llama-3.2-3B-Instruct, Granite-4.1-8B→Phi-4-mini-instruct, Granite-4.1-8B→Qwen2.5-7B-Base
- Data: 20,000 prompts (10,000 math from DAPO-Math-17K, 10,000 code from CodeForces)
- Evaluation: MATH500, GSM8K, AIME-2024/2025/2026, AMC23, Minerva-Math (math); HumanEval, MBPP, LiveCodeBench (code)
- Metrics: mean@32 for math, mean@8 for code
Gradient Diagnostics
Directional agreement (cosine) and relative magnitude (norm ratio) between strict and span gradients:
Empirical Validation / Results
1. Strict Alignment Coverage vs. Static Vocabulary Overlap
Table 1: Static vocabulary Jaccard overlap and strict token coverage on student trajectories
| Teacher→Student | Coverage | Steps 1-100 | Steps 1-20 | Steps 41-60 | Steps 81-100 | |
|---|---|---|---|---|---|---|
| Qwen→Llama | 64.32 | 93.56 | 94.63 | 93.20 | 93.43 | |
| 85.91 | 88.65 | 84.90 | 85.36 | |||
| Granite→Phi | 39.49 | 96.98 | 96.65 | 96.95 | 97.12 | |
| 97.26 | 95.83 | 97.50 | 97.96 | |||
| Granite→Qwen | 64.87 | 85.57 | 85.23 | 85.06 | 85.88 | |
| 82.82 | 81.15 | 82.77 | 83.86 |
Key insight: Granite→Phi has the lowest Jaccard overlap (39.49%) but the highest strict token coverage (>96%). Substantial static vocabulary mismatch can coexist with strict alignment at most positions on student-generated trajectories.
2. Span Supervision Reduces Accuracy
All three pairs attain their highest full average at (strict only). The 18 positive-weight settings score 0.27–1.20 percentage points below their respective strict baselines, with an overall downward trend as λ increases.
3. Probability Mass Concentrates on Shared Vocabulary
The shared vocabulary retains 99.69–99.90% of teacher mass and 98.99–99.81% of student mass on average. At k=16, the student-selected top-k subset retains at least 93.54% of teacher mass and 94.55% of student mass.
4. Top-k Subset Training Results
Table 2: Cross-tokenizer distillation accuracy (%) — Bold marks the best distilled score per column
| Method | Qwen→Llama Math | Qwen→Llama Code | Qwen→Llama Full | Granite→Phi Math | Granite→Phi Code | Granite→Phi Full | Granite→Qwen Math | Granite→Qwen Code | Granite→Qwen Full |
|---|---|---|---|---|---|---|---|---|---|
| Base | 18.32 | 35.60 | 26.96 | 33.06 | 40.04 | 36.55 | 24.61 | 34.51 | 29.56 |
| ULD | 20.51 | 37.74 | 29.12 | 35.24 | 42.01 | 38.63 | 33.36 | 20.87 | 27.12 |
| Extended ULD | 23.01 | 39.26 | 31.13 | 35.28 | 41.42 | 38.35 | 33.14 | 32.00 | 32.57 |
| GOLD | 21.21 | 33.10 | 27.16 | 36.22 | 47.29 | 41.75 | 38.26 | 54.84 | 46.55 |
| SimCT | 25.31 | 37.86 | 31.59 | 36.67 | 46.70 | 41.68 | 38.64 | 53.56 | 46.10 |
| Strict full | 26.60 | 39.11 | 32.86 | 37.49 | 47.49 | 42.49 | 39.85 | 54.84 | 47.35 |
| Strict top-16 | 26.86 | 38.42 | 32.64 | 37.23 | 47.29 | 42.26 | 39.07 | 55.06 | 47.06 |
| Strict top-128 | 26.31 | 38.44 | 32.38 | 37.27 | 47.82 | 42.54 | 39.67 | 55.62 | 47.65 |
Key findings:
- All three strict variants exceed all four alternatives in both math and full averages on every model pair
- Top-16 retains at least 96% of the full-average improvement from Base to strict full
- Top-16 is ahead of the strongest alternative in the full average by 0.51–1.05 percentage points
- No consistent further gain from increasing subset size beyond 16
5. Gradient Diagnostics
- Directional agreement: Mismatch cosines are near zero for Qwen→Llama and Granite→Phi; Granite→Qwen starts with negative agreement approaching zero later. The split-half strict cosine remains higher than the mismatch–strict cosine at every checkpoint.
- Relative scale: The span-to-strict gradient norm ratio increases across measured checkpoints in all three pairs, meaning a fixed positive weight on the span loss gives it increasing gradient norm relative to the strict component.
Theoretical and Practical Implications
Theoretical Implications
-
Coverage ≠ Learning Value: The paper provides systematic evidence that expanding structural supervision coverage does not guarantee better distillation. This challenges the implicit assumption in prior work (e.g., SimCT) that recovering excluded supervision signals is beneficial.
-
Support Concentration Under Heterogeneous Tokenizers: The finding that a compact top-16 subset retains most distillation gains extends prior work on support concentration (Fu et al., 2026; Li et al., 2026) to the cross-tokenizer setting, showing that student-selected supports remain effective even when vocabularies only partially overlap.
-
Gradient Interaction as Diagnostic Tool: The gradient cosine and norm ratio analyses provide a mechanistic explanation for why span supervision hurts: weakly aligned gradients with growing relative magnitude can interfere with the strict objective.
Practical Implications
-
Efficient Cross-Tokenizer Distillation: Practitioners can restrict distillation to strict 1:1 positions with a top-16 shared vocabulary subset, achieving near-full performance with substantially lower computational cost.
-
Supervision Quality Assessment: The paper suggests that before adding new supervision signals, practitioners should evaluate their directional agreement with existing losses and their relative gradient scale.
-
Baseline Comparisons: Strict full and Strict top-16 provide strong baselines that outperform more complex cross-tokenizer methods (ULD, GOLD, SimCT), suggesting that simpler approaches may be more effective.
Conclusion
The paper re-examines the role of alignment coverage in Cross-Tokenizer OPD across three heterogeneous teacher–student pairs on mathematical reasoning and code generation. The main takeaways are:
-
Strict alignment is already substantial: Strict 1:1 groups cover most student-generated tokens despite large static vocabulary mismatch.
-
Shared vocabulary retains most probability mass: On student-generated responses, the shared vocabulary carries nearly all teacher and student predictive probability at strict positions.
-
Compact supervision is effective: Restricting reverse KL to a student-selected top-16 subset retains at least 96% of full shared-vocabulary OPD improvement, outperforming all evaluated cross-tokenizer baselines.
-
Span supervision is harmful: Adding span log-probability MSE achieves complete supervision coverage but reduces accuracy. Gradient diagnostics show span gradients have weak/negative directional agreement with strict gradients and grow in relative magnitude during training.
Future directions suggested by this work include:
- Developing principled methods for assessing supervision reliability before adding new training signals
- Investigating whether other forms of mismatch-group supervision (beyond MSE on log-probabilities) could be more effective
- Extending the analysis to other domains and larger model scales (the paper includes preliminary results with a 235B teacher in the appendix)
Related papers
- From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining
Search scaling in autonomous research is capability-dependent: deeper search narrows cross-model gaps, but early trajectory state and parallel exploration critically shape final performance.
- On-Demand Attention: Language Models Know When to Recall
On-demand attention uses a lightweight recall head to predict when global attention helps, recovering most quality with up to 2.65x decoding throughput.
- Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
Identical runs of the same AI coding agent vary more than differences between agent-model pairings, so best-of-three with compliance checking beats single-run benchmarking.