Full text not available for this paper
Summary (Overview)
- Central claim: Routing drift alone is insufficient evidence of routing failure in merged Mixture-of-Experts (MoE) LLMs; source-informed corrections must be judged by their task-level intervention effects, not by structural routing disagreement.
- Key finding 1: Across DeepSeekMoE, OLMoE, and Qwen3-MoE under Average and Task Arithmetic merging, 77.7–96.9% of changed-route events are attributable to input (representation) shifts rather than router parameter changes.
- Key finding 2: Structural routing differences (e.g., JS divergence) poorly predict next-token likelihood gains from source-route restoration (AUROC 0.47–0.52, near chance), and different expert selections can produce directionally similar mixture outputs (cosine similarity 0.896–0.976).
- Key finding 3: The paper formalizes routing failure as intervention-relative recoverable task loss with non-routing parameters fixed, and demonstrates that controlled router corruption is recoverable (12.50–32.81 pp) while source-route restoration yields no reliable task benefit.
- Key finding 4: The proposed Selective Router Repair (SRR) method—which fits likelihood-weighted source-derived expert-pair corrections—shows no reliable evidence that source-likelihood advantages identify beneficial corrections or improve average task performance (−0.056 to +0.139 pp across settings).
Introduction and Theoretical Foundation
Background: Model merging combines specialized LLMs without joint retraining. For MoE models, merging changes not only parameters but also token-to-expert routing—termed routing drift. Recent methods (e.g., HARC) treat routing mismatch as "routing breakdown" and realign merged routers.
Core distinction: The paper distinguishes two concepts:
- Routing drift: what changed in expert assignment and routing probabilities
- Routing failure: degradation in task-relevant behavior attributable to routing
Theoretical basis: A token's route depends jointly on (1) its router-input representation (hidden state), (2) the router parameters, and (3) the weighted outputs of selected experts. Therefore, routing drift alone reveals neither why routing changed nor whether it caused harm. The MoE layer output is:
where is the selected expert set, the routing weight of expert , and its output at layer .
Research questions: (1) Does routing drift after MoE merging actually indicate routing failure? (2) What evidence should justify repair?
Methodology
Models and merging methods: DeepSeekMoE-16B-A3B (Top-6 routing, 27 sparse layers), OLMoE-7B-A1B (Top-8, 16 layers), Qwen3-30B-A3B (Top-8, 48 layers), each evaluated under Average, Task Arithmetic (TA), TIES, and WUDI-Merge.
Crossed intervention analysis: For source and merged quantities, four states are evaluated:
where is the expert selection, denotes router parameters, and the router input. Conditioning on : a change is representation-induced when but ; router-parameter-induced is the converse.
Task-grounded diagnosis: Define intervention-relative recoverable loss:
where and are baseline and alternative routing policies, are fixed non-routing parameters. Positive values indicate recoverable loss.
Selective Router Repair (SRR): Constructs source–base preference profiles per prompt:
where are router logits, are NLLs, is sigmoid, and center subtracts the expert-wise mean. For each selected pair , the fitting objective is:
with . Router rows update as and .
Evaluation: 8 benchmarks (MMLU, HellaSwag, ARC-C, ARC-E, PIQA, WinoGrande, BoolQ, GSM8K) with 35,326 matched items per pair, five evaluation runs, paired item-bootstrap 95% confidence intervals.
Empirical Validation / Results
Routing origin attribution (Table A2, Figure 2):
| Architecture / Parent | Changed route (%) | Representation (%) | Router parameters (%) |
|---|---|---|---|
| DeepSeekMoE / Average | 33.9 [32.2, 35.7] | 77.7 [77.1, 78.2] | 1.79 [1.65, 1.94] |
| DeepSeekMoE / TA | 46.7 [44.9, 48.7] | 78.5 [78.1, 79.0] | 1.14 [1.04, 1.23] |
| OLMoE / Average | 42.7 [41.3, 44.1] | 85.8 [85.4, 86.1] | 0.62 [0.57, 0.68] |
| OLMoE / TA | 54.3 [52.9, 55.7] | 87.0 [86.6, 87.4] | 0.36 [0.33, 0.40] |
| Qwen3-MoE / Average | 26.46 | 96.87 | 0.034 |
| Qwen3-MoE / TA | 33.20 | 96.90 | 0.040 |
Structural predictors near chance: JS divergence predicts source-route NLL gain with AUROC 0.47–0.52 (chance = 0.5). None of the 16 primary tests survives Holm correction.
Mixture output similarity (Figure 4): Source-route and native merged-route mixtures have mean cosine 0.896–0.976, while expert-pair maximum cosine is only 0.041–0.096. Observed routes exceed matched-random controls by 0.112–0.147 in cosine.
Controlled corruption recovery (Figure 5): Permuting router logits in OLMoE layers recovers 12.50–14.84 pp (4 layers) and 26.56–32.81 pp (16 layers) via clean-route replay. All eight comparisons at k=4, 16 pass Holm correction.
Natural routing alternatives (Table 1): Source-route replay and LC replay yield no strictly positive intervals for accuracy or margin changes across all settings.
SRR task results (Table 2): Average score changes range from −0.056 to +0.139 pp across eight OLMoE/Qwen3-MoE settings. All 95% CIs include zero. HARC and SRR each lead in four of eight settings.
Local utility diagnostics: Source and fitted directions agree in sign on 83.4% of supported events, yet neither pooled (gain over native) nor (directional gain) establishes improvement. Matched selection effects show no enrichment (Table 3):
| Contrast | Estimate | 95% CI |
|---|---|---|
| Native () | −0.341 | [−3.323, 2.631] |
| Opposite () | −0.997 | [−5.653, 3.868] |
Theoretical and Practical Implications
Theoretical implications:
- Routing drift is a structural phenomenon; routing failure is a causal, task-relative phenomenon. These must not be conflated.
- The paper provides a formal framework (intervention-relative recoverable loss) for diagnosing routing failure that separates candidate construction from demonstrated recovery.
- Different routes can preserve similar mixture outputs (functional redundancy), explaining why routing disagreement need not imply harm.
Practical implications:
- Post-merge routing "repair" methods (e.g., HARC) should be evaluated against task-level intervention effects, not source-route agreement.
- Source-likelihood advantages do not reliably identify beneficial router corrections; repair decisions require task-grounded validation.
- The released analysis toolkit enables controlled counterfactual interventions and paired token-/task-level evaluation for future MoE merging research.
Conclusion
The paper demonstrates that routing drift alone is insufficient evidence of routing failure in merged MoE LLMs. Most post-merge route changes are representation-induced rather than router-parameter-induced, structural differences poorly predict intervention gains, and different routes can preserve similar mixture outputs. The formalization of routing failure as intervention-relative recoverable task loss—with non-routing parameters fixed—provides a principled diagnostic framework. Controlled corruption tests confirm recoverability under deliberate router corruption, while source-route restoration shows no reliable task benefit. The proposed SRR case study finds no evidence that source-likelihood advantages identify beneficial corrections or that fitted updates improve average task performance.
Future directions: (1) Exploring alternative routing interventions beyond source-route restoration that might yield recoverable task loss; (2) developing selection criteria for router repair grounded in task-level intervention effects rather than source agreement; (3) extending the analysis to additional MoE architectures and merging methods; (4) investigating whether other supervision signals (beyond source-likelihood) can identify beneficial local corrections.
Related papers
- Counterexamples to Local Reconstruction Gain as a Proxy for Final Fidelity in Residual Completion
Local attention-output reconstruction gains do not guarantee final-model fidelity, as residual completion can improve local error while worsening dense-model KL divergence.
- Cost-free Spectral Estimation for Adaptive Newton--Schulz in Matrix Optimizers
Newton–Schulz iterations expose spectral moments via free scalar reductions, enabling spectrum-adaptive routine selection that cuts polar error up to 90x and improves LLM pretraining loss.
- Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons
Fast learning-rate transfer holds at growing training horizons when T grows slower than sqrt(n), but requires nondegenerate first-order loss sensitivity to avoid spectral-dependent failures.