MoE & Sparse ExpertsIssue 9Oct 3 – 10, 2026

Routing Consistency: From Replay to Semantic Anchoring and Robustness Objectives

Highlights of This Issue

The main theme of this issue is that "routing consistency is moving from replay to semantic anchoring and robustness objectives," while load balancing is being repositioned on the service side as a constraint rather than a goal.

On the training-inference routing consistency side, Structuring MoE Expert Selection for Agentic RL for the first time uses turn-level operation labels (such as READ/UPDATE) from agentic trajectories as explicit supervision signals for MoE routing specialization. It encourages turns with the same operation to share experts and turns with different operations to separate experts via mutual information maximization, complemented by selective token-level local consistency regularization and entropy gating for stable training. Unlike Terminal Agent RL from Issue 8 (R3 records and replays actual routing) and ESRL from Issue 6 (entropy-adaptive perturbation + replay), this paper does not replay routing but directly reshapes the routing distribution anchored by operation semantics, offering an alternative path beyond routing replay. It reports complete routing statistics and dense baselines on Qwen3-30B-A3B and Qwen3.5-35B-A3B. DRMoET approaches from the training objective side: even with perfectly balanced routing, insufficient capacity of non-top experts can still lead to misrouting losses. Therefore, it treats each layer's experts as endogenous robustness groups, using EMA-smoothed activation-weighted expert losses to update dual weights, directly optimizing high-loss routing outcomes rather than merely balancing traffic. On a 10.3B scale, it improves the seven-task average from 0.6625 to 0.6767 and provides mechanism analyses such as forced misrouting probes and expert loss variance.

On the load balancing side, Zepp for the first time downgrades load balancing from an optimization objective to a physical resource (GPU/NIC) constraint, directly optimizing the bottleneck cross-node communication. It proposes split/merge bidirectional communication shaping primitives and intra-node expert exchange (avoiding cross-node weight migration), reporting comparisons with 7 baselines and weak scaling experiments on 4-16 nodes. CIPHER-MoE combines expert-side capacity control with similarity ranking in trillion-parameter training, providing a sufficient bound for the expert capacity required by similarity selection. On DeepSeek-V4-Pro 1.6T, it validates a 64.9 percentage point reduction in Top-1 expert load and 1.10–1.94× training speedup.

On the expert specialization side, RouterInterp proposes and empirically tests the "Superposition Specialization Hypothesis" (SSH): experts are not specialized in a single semantic domain but in the union of multiple semantically unrelated micro-domains. Using SAE features as proxies for micro-domains, it quantifies the multi-semantic nature of expert routing, achieving an explanation F1 of 0.49 on gpt-oss-20b, significantly outperforming token statistics baselines. On the system side, Expert Coupling in MoE Pretraining for the first time uses expert co-selection correlations for both intra-layer placement and cross-layer token rearrangement during pretraining without altering routing decisions; RelayMoE proposes a ring-based MoE execution model that cyclically passes expert weights or token activations to avoid constructing full top-k expansion distribution buffers; DivMoE attributes the collapse pathology of fine-grained upcycling to the combination of single-source initialization and fine-grained splitting, proposing domain-specialized initialization and hard diversity-constrained routing.

Community and Updates

A practical record on MoE optimization peculiarities reports a candid negative result: Router Policy Optimization fixes for router/target mismatch did not improve pretraining on GPT-2/TinyStories. The author acknowledges that the fix "did not improve large-scale MoE pretraining"—providing a counterexample baseline to the common assumption that fixing routing mismatch necessarily yields gains. Understanding, Extending, and Improving Quantile Balancing points out that allowing slack around uniform load can improve loss, directly challenging the default strict uniformity assumption in auxiliary-loss-free methods.

Open Questions

  1. Structuring MoE Expert Selection anchors routing with operation semantics, while Terminal Agent RL from Issue 8 replays actual routing with R3—can semantic anchoring and routing replay be combined? How do the granularities of operation labels (turn-level vs. token-level) complement the fidelity of routing replay?

  2. How can the robustness objective of DRMoET be unified with the ID Balancing load control from Issue 8 and the quantile bias of Kimi K3 from Issue 1 in terms of training objectives? Does robustness optimization conflict with load balancing goals, or can they be stacked?

  3. Zepp downgrades balancing to a constraint, and CIPHER-MoE performs capacity control on the training side—how can the slack boundary of "balancing as a constraint" be quantitatively characterized? Under what load distributions is it worthwhile to abandon strict balancing for communication optimization?

  4. What does the Superposition Specialization Hypothesis of RouterInterp imply for the causal audit of expert importance in From Observation to Intervention from Issue 1? If experts specialize in unions of micro-domains, does the burden of proof for single-expert causal intervention need to be redefined?

Papers in this issue

  1. Hierarchical routing control that aligns MoE expert selection with agentic trajectory structure improves success rates by over 10 points and boosts inference throughput up to 41.1% without architectural changes.

    Editor's note

    Must-read. For the first time, turn-level operation labels (READ/UPDATE, etc.) from agentic trajectories are used as explicit supervision signals for MoE routing specialization, encouraging turns with the same operation to share experts and different operations to separate experts via mutual information maximization, with entropy gating for stable training. Compared to Terminal Agent RL from Issue 8 (R3 records and replays actual routing) and ESRL from Issue 6 (entropy-adaptive perturbation + replay), this paper does not replay routing but directly reshapes the routing distribution anchored by operation semantics, offering an alternative path beyond routing replay. It reports complete routing statistics and dense baselines on Qwen3-30B-A3B and Qwen3.5-35B-A3B.

  2. Zepp redefines load balancing as a constraint in distributed MoE serving, achieving up to 6.68x speedup by directly minimizing inter-node communication through avoid, reshape, and hide strategies.

    Editor's note

    Must-read. For the first time, load balancing is downgraded from an optimization objective to a physical resource (GPU/NIC) constraint, directly optimizing the bottleneck cross-node communication, and proposing split/merge bidirectional communication shaping primitives and intra-node expert exchange (avoiding cross-node weight migration). Compared to MegaFlux pipelined replication from Issue 8 and PipelinedLLEP memory-side mitigation from Issue 6, this paper explicitly distinguishes GPU and NIC constraints and centers on communication shaping, reporting comparisons with 7 baselines and weak scaling experiments on 4-16 nodes.

  3. DRMoET, a distributionally robust MoE training objective, improves downstream task accuracy by up to 2.1% over baselines by optimizing worst-case routing outcomes rather than just load balancing.

    Editor's note

    Must-read. For the first time, the robust optimization idea of Group DRO is introduced into MoE training, treating each layer's experts as endogenous robustness groups, using EMA-smoothed activation-weighted expert losses to update dual weights, directly optimizing high-loss routing outcomes rather than merely balancing traffic. Compared to the ID Balancing control-theoretic framework from Issue 8 and Kimi K3 Quantile Balancing from Issue 1, this paper shifts from load balancing to expert capability robustness, representing a new path at the training objective level, reporting mechanism analyses such as forced misrouting probes and expert loss variance on a 10.3B scale.

  4. CIPHER-MoE provably reduces expert capacity needs via affinity-aware token selection, achieving 1.10x–1.94x training speedups on trillion-parameter MoE LLMs while preventing OOM failures.

    Editor's note

    Worth reading. For the first time, expert-side capacity control is combined with similarity ranking in trillion-parameter MoE training, proposing bidirectional token-expert assignment and providing a sufficient bound for the expert capacity required by similarity selection. Compared to ACE inference-time expert skipping from Issue 6 and Training-Free Halving inference-time k reduction from Issue 5, this paper focuses on the training phase, controlling hotspot expert load through expert-side similarity filtering and rerouting, validating training speedup and quality preservation on DeepSeek-V4-Pro 1.6T.

  5. RouterInterp shows MoE experts specialize in disjoint collections of unrelated micro-domains, not coherent semantic domains, challenging the standard specialisation hypothesis.

    Editor's note

    Worth reading. For the first time, the "Superposition Specialization Hypothesis" (SSH) is proposed and empirically tested: experts specialize in the union of multiple semantically unrelated micro-domains rather than a single semantic domain, using SAE features as proxies for micro-domains to quantify the multi-semantic nature of expert routing. Compared to the routing effective rank dynamics diagnosis from Issue 5 (From Concentration to Differentiation and Back) and the causal audit of expert importance from Issue 1 (From Observation to Intervention), this paper is the first to use SAE features as micro-domain proxies, achieving an explanation F1 of 0.49 on gpt-oss-20b, providing a new quantitative dimension for evidence standards of expert specialization.

  6. Expert routing in MoE models is highly predictable, enabling correlated expert placement and token shuffling that cut all-to-all communication time by up to 2.63x without altering training.

    Editor's note

    Worth reading. For the first time, expert co-selection correlations are used for both intra-layer placement and cross-layer token rearrangement during pretraining: intra-layer uses co-selection graphs for Kernighan-Lin local search placement with a deduplication scheduler, and cross-layer uses the first two layers' routing to predict the GPU for the next layer's experts, folding token migration into the sequence-parallel reduce-scatter. Compared to Cobalt's node-level co-activation communication reduction from Issue 8, this paper advances co-selection signals to intra-layer placement and cross-layer token migration without altering routing decisions, reducing all-to-all time by 1.16–2.63× on Megatron-LM.

  7. RelayMoE replaces all-to-all MoE dispatch with ring-based local computation, cutting peak memory up to 7.6x and speeding training throughput by 2x.

    Editor's note

    Worth reading. Proposes a ring-based MoE execution model that cyclically passes expert weights or token activations to avoid constructing full top-k expansion distribution buffers, significantly reducing peak memory, and supports adaptive selection between expert routing and token routing. Compared to PipelinedLLEP from Issue 6, which mitigates routing skew costs from the memory side, this paper changes data flow organization from the execution side, reducing peak memory by 7.6× and increasing trainable sequence length by 2.85× on 30B–57B models.

  8. DivMoE achieves fine-grained MoE upcycling by combining domain-specialized experts with diversity-constrained routing, preventing routing collapse and outperforming all baselines across 15 benchmarks.

    Editor's note

    Worth reading. For the first time, the collapse pathology of fine-grained MoE upcycling is attributed to the combination of single-source initialization and fine-grained splitting, proposing domain-specialized initialization and hard diversity-constrained routing (at most one expert per domain group per token) as a joint solution. Compared to the sparsity-granularity decoupling of Slicing and Dicing from Issue 3, this paper adds domain diversity as a new axis, improving fine-grained upcycling from 23.2% to 55.6% on Qwen3-1.7B, providing a new empirical baseline for fine-grained upcycling.