RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
Summary (Overview)
-
Core contribution: The paper introduces the Superposed Specialisation Hypothesis (SSH), which posits that Mixture-of-Experts (MoE) models' experts specialize in disjoint collections of unrelated micro-domains rather than single coherent semantic domains, challenging the commonly assumed Domain Specialisation Hypothesis (DSH).
-
Novel method: The authors propose RouterInterp, an SAE-based interpretability method that identifies sparse autoencoder features most predictive of routing decisions and generates unified natural language explanations of expert behavior.
-
Key empirical finding: Expert routing is highly polysemantic—each expert's most predictive SAE latents span ~11 semantically distinct clusters (out of 20), far from the single cluster predicted by DSH.
-
Performance results: RouterInterp achieves explanation scores of 0.49 on gpt-oss-20b and 0.60 on OLMoE-1B-7B, outperforming token-statistics-based methods by ~65% and prior expert-impact methods by ~28% on gpt-oss-20b.
-
Theoretical contribution: The paper provides two arguments for why SSH holds: interference minimization (analogous to superposition theory) and load-balancing pressure during training.
Introduction and Theoretical Foundation
Background
Sparse Mixture-of-Experts (MoE) transformers scale more efficiently than dense models by routing tokens to modular expert networks, activating only 2–15% of parameters per input. The strong performance of MoE models has often been attributed to expert specialisation—the idea that each expert learns to handle a subset of the input distribution.
However, prior interpretability efforts have struggled to recover clear, interpretable patterns of expert specialisation. The paper identifies a crucial distinction between two forms of the Specialisation Hypothesis:
Domain Specialisation Hypothesis (DSH): Experts specialize in different semantically coherent domains (e.g., medical text, mathematical reasoning, romance language translation). Semantically similar micro-domains cluster within an expert.
Superposed Specialisation Hypothesis (SSH): Experts specialize in disjoint collections of features spanning multiple unrelated micro-domains. This mirrors the superposition phenomenon observed in neural networks.
Theoretical Foundation
The key insight is the Pigeonhole Principle argument: real-world data contains far more fine-grained categories (micro-domains) than any MoE layer has experts. When (more micro-domains than experts), multiple micro-domains must inevitably be routed to the same expert. The question becomes whether those micro-domains are similar (DSH) or dissimilar (SSH).
Motivation for SSH
Argument 1 — Interference Minimisation: Drawing on superposition theory (Elhage et al., 2022), the authors argue that semantically distinct micro-domains that rarely co-activate can share an expert with minimal interference noise, enabling "Computation in Superposition" (Hänni et al., 2024).
Argument 2 — Load Balancing: Load-balancing losses encourage uniform expert utilization. With small batch sizes, a batch contains few macro-domains, but the router must distribute tokens from these few domains across all experts—necessarily routing dissimilar inputs to the same expert.
Methodology
Sparse Autoencoders (SAEs)
The paper uses Top-K Sparse Autoencoders to decompose model activations into interpretable features. The SAE architecture maps an activation vector to a sparse latent representation (where ):
where , , and retains only the largest activations.
MoE Routing
In a standard MoE layer, a router network computes routing logits and converts them to a probability distribution over experts via softmax:
Only the top- experts are selected, and the layer output is:
RouterInterp Method
RouterInterp operates in three stages:
-
Identifying Features: For each expert , select the top- SAE features most predictive of routing by approximating the effect of ablating each feature on expert selection. Each feature is assigned to at most one expert.
-
Collecting Activations: For each selected feature , collect context windows where is active, forming:
- Positive set : active AND token routed to
- Negative set : active but token NOT routed to
-
Explaining: Prompt a language model with positive-negative contrast groups per feature to synthesize a unified natural language description of when expert is selected.
Evaluation Setup
- Models: OLMoE-1B-7B (64 experts, ) and gpt-oss-20b (32 experts, )
- SAEs: Top-K SAEs trained on 100M tokens (OLMoE) with 32,768 features, sparsity ; BatchTopK SAEs for gpt-oss with 131,072 features, sparsity
- Scoring: Adapted AutoInterp Detection scoring—a scorer LLM uses the explanation to predict expert activation on held-out examples, with F1 as the metric
- Baselines: Unigram Lookup (token co-occurrence), Expert Impact AutoInterp (Herbst et al., 2026)
Empirical Validation / Results
Evidence for SSH
The paper tests whether each expert's top-20 predictive SAE latents are mutually similar (DSH) or dissimilar (SSH). They define as the number of semantically distinct clusters that expert 's top- latents fall into:
- DSH predicts: (single coherent domain)
- SSH predicts: (each latent in its own cluster)
- Observed: Mean = 11.3 on gpt-oss-20b (layer 20) and 10.8 on OLMoE-1B-7B (layer 15), out of upper bound
The observed values lie far above the DSH prediction of , providing strong evidence for SSH.
RouterInterp Performance
Figure 4 results (mean F1 scores):
| Method | gpt-oss-20b | OLMoE-1B-7B |
|---|---|---|
| RouterInterp | 0.492 (s=128) | 0.602 |
| Expert Impact AutoInterp | 0.383 | 0.305 |
| Unigram Lookup | 0.299 | 0.342 |
Ablation Study (Table 1, gpt-oss-20b)
| Method | Layer 4 | Layer 8 | Layer 12 | Layer 16 | Layer 20 |
|---|---|---|---|---|---|
| Unigram Lookup | 0.295 | 0.300 | 0.278 | 0.305 | 0.315 |
| Bigram Lookup | 0.309 | 0.320 | 0.283 | 0.320 | 0.331 |
| Unigram AutoInterp | 0.332 | 0.406 | 0.320 | 0.356 | 0.348 |
| Bigram AutoInterp | 0.429 | 0.453 | 0.385 | 0.407 | 0.411 |
| Expert Activations AutoInterp | 0.415 | 0.475 | 0.345 | 0.425 | 0.425 |
| RouterInterp (s=64) | 0.490 | 0.542 | 0.465 | 0.509 | 0.470 |
| RouterInterp (s=128) | 0.531 | 0.509 | 0.456 | 0.516 | 0.448 |
Key Ablation Insights
- Verbalizing token statistics with an LLM improves explanations only modestly (0.30→0.35 mean F1)
- Replacing token lists with full activating passages does not help (Expert Activations AutoInterp ≈ Bigram AutoInterp at 0.417 mean F1)
- Bigram descriptions are narrow (higher precision 0.39, lower recall 0.59); full-window descriptions are broad (lower precision 0.30, higher recall 0.81)
- Only when evidence is grouped by SAE feature do precision and recall improve together (P 0.47 / R 0.64; 0.495 mean F1, s=64)
Theoretical and Practical Implications
Implications for Interpretability
-
Explains prior failures: The SSH explains why previous work assuming monosemantic experts (implicitly following DSH) was unable to successfully interpret routing behavior.
-
Reconciles apparent specialisation: Cases of apparent domain specialisation (e.g., code generation experts, language-specific experts, safety refusal experts) are consistent with SSH—an expert may participate in a task without that being its only function.
-
SAE-based discovery: The success of RouterInterp supports the view that SAEs are useful for "discovery of unknowns" rather than acting on known concepts (Peng et al., 2026).
Theoretical Implications
-
Computation in Superposition: Experts may group seemingly unrelated inputs that require a set of disjoint computational operations, suggesting a functional similarity (rather than input similarity) basis for routing.
-
Design implications: Understanding whether experts cluster by input similarity or functional similarity could inform the design of more efficient and interpretable MoE architectures.
-
Scaling insights: With much larger numbers of experts, routing becomes increasingly monosemantic and interpretable (Park et al., 2025), suggesting a trade-off between interpretability and compute efficiency.
Practical Implications
- RouterInterp provides a scalable method for generating accurate explanations of expert routing, potentially enabling:
- Better debugging and auditing of MoE models
- More targeted interventions in expert behavior
- Improved transparency for safety-critical applications of foundation models
Conclusion
The paper makes three main contributions:
-
Theoretical: Introduces and provides evidence for the Superposed Specialisation Hypothesis, formalizing the distinction between domain-level and feature-level specialisation in MoE models.
-
Empirical: Demonstrates that expert routing is highly polysemantic, with experts' predictive features spanning many semantically distinct clusters.
-
Methodological: Presents RouterInterp, which achieves state-of-the-art explanation scores (0.49 on gpt-oss-20b, 0.60 on OLMoE-1B-7B) by leveraging SAE features to enumerate experts' disjoint micro-domains.
Future Directions
- Mechanisms of specialisation: Why do experts converge to particular feature combinations? Does load-balancing force redundancy?
- Functional vs. input similarity: Do co-routed tokens share semantic content or require similar downstream transformations?
- Developmental interpretability: How does routing evolve across training checkpoints?
- Scaling: Apply RouterInterp to frontier architectures with more parameters, shared experts, and different routing methods (e.g., Expert Choice).
Limitations
- Dependence on SAE quality (performance correlates with reconstruction quality/FVU)
- Explanations grow in size with more features; compression and interactive formats (e.g., Neuronpedia dashboards) could help
- Evaluated only on 1B–20B parameter models; scaling to frontier architectures remains open
Related papers
- Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position
Hybrid position combining NoPE with position-biased attention, not hybrid architecture alone, drives long-context performance, with SWLA enabling 16x training-free length extrapolation.
- The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
Self-improving LLM loops on small reused evaluation sets inflate reported gains by 13–20 points, with most proposals after the first rewrite being harmful.
- How to scale your HEP ML models: A recipe for robust architecture comparisons at scale
A hyperparameter recipe makes learning-rate and batch-size scaling predictable, yielding Chinchilla-like sqrt(C) compute-optimal scaling for jet-tagging transformers on ATLAS data.