MoE & Sparse ExpertsIssue 8Sep 26 – Oct 3, 2026

Load Balancing Enters the Era of Cybernetics and Precision, Evidence Standards for Routing Drift Are Re-examined

Highlights of This Issue

The main theme of this issue is that "load balancing is moving from heuristic approaches to cybernetics and precision," while the evidence standards for routing drift are being directly challenged in the context of model merging.

On the load balancing side, ID Balancing for the first time unifies auxiliary-loss-free load balancing methods into a PID control framework: DeepSeek's loss-free method is identified as a fixed-step integral controller, and the Quantile Balancing from Issue 1's Kimi K3 is identified as a generalized proportional controller. Based on this, the paper proposes ID Balancing with a magnitude-aware integral term and a deterioration-gated derivative term, making corrections adaptive to error magnitude and deterioration direction, and validates stability advantages on Top-3/5/10-of-768 extreme sparsity and 69.9B scale. Exact Quantile Balancing advances from another angle: it uses two-pass BF16 radix selection to recover the global batch empirical quantiles, with communication cost independent of token count and invariant to data partitioning, solving the partition dependence caused by global quantile approximation in distributed Quantile Balancing; its Load-Error Injection uses the constant diagonal Jacobian of unnormalized router scores as an STE proxy to inject local load errors, avoiding the cross-expert coupling caused by normalized probabilities in the GShard loss. These two works respectively advance the Quantile Balancing from Issue 1's Kimi K3 into provably stable and precise solutions from the perspectives of cybernetics and precision.

On the routing drift side, this issue presents a pair of tensions. Terminal Agent RL for the first time decomposes training-inference consistency into two orthogonal axes—token fidelity (TITO) and routing fidelity (R3)—in agentic RL at 122B scale MoE. R3 records and replays the actual expert selections per layer during rollout, reducing the log-prob gap from 0.021 to 0.013 on long-horizon tasks with 300+ tool calls—this is the first validation of routing replay (predictive replay from Issue 1's PR²) at frontier scale and in agentic scenarios. Meanwhile, Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs systematically falsifies the notion that "routing drift implies routing failure" in model merging scenarios: cross-interventions show that most expert reassignments stem from representation shifts rather than router parameter changes, and structural distance cannot predict source routing recovery gains (AUROC 0.47–0.52). Reading both together suggests: whether routing drift is worth fixing depends on whether task-level intervention benefit evidence can be provided.

On the systems side, MegaFlux for the first time directly pipelines the weight transfer and gradient reduction caused by dynamic expert replication into the persistent scheduling of MoE megakernels, using phase-level weight-ready gating for forward, gradient tile priority production, and reduction-overlapped backward, pipelining replication costs from the execution side (contrasting with the memory-side mitigation in Issue 6's PipelinedLLEP); Cobalt uses "expert co-activation" as an explicit signal for cross-node communication reduction, using Ochiai similarity to advance layout planning from single-expert statistics to expert-pair statistics.

Open Questions

  1. Can the PID perspective of ID Balancing and the exact quantiles of Exact Quantile Balancing be unified? Can the cybernetic framework absorb exact quantile computation to simultaneously achieve stability and partition invariance?

  2. The R3 replay of actual routing in Terminal Agent RL brings benefits, while Routing Drift Alone proves that drift itself does not diagnose failure—under what conditions is routing drift worth fixing? Can task-level intervention benefit serve as a unified criterion?

  3. The LEI in Exact Quantile Balancing uses an STE proxy to inject local load errors. How do its gradient properties quantitatively compare with GShard-type auxiliary losses and the quantile bias from Issue 1's Kimi K3 in terms of cross-expert coupling and stability?

  4. The expert-pair co-activation signal in Cobalt is effective in training layouts. Can this signal transfer to the serving side (e.g., the online lookahead placement in Issue 4's Director) to reduce cross-node communication during inference?

Papers in this issue

  1. ID Balancing, a magnitude-aware integral-derivative controller with worsening-gated derivative and bias re-centering, reduces worst-case expert overload by over 50% versus best baselines in sparse MoE training.

    Editor's note

    Must-read. For the first time unifies auxiliary-loss-free load balancing into a PID control framework: DeepSeek's loss-free is fixed-step integral control, Issue 1's Kimi K3 Quantile Balancing is generalized proportional control. Based on this, the paper proposes ID Balancing with magnitude-aware integral term and deterioration-gated derivative term, validating stability advantages on Top-3/5/10-of-768 extreme sparsity and 69.9B scale, reporting full routing statistics and dense baselines. Compared to Issue 1's Kimi K3, this paper provides a cybernetic unified perspective and provable stability increments.

  2. Exact Quantile Balancing and Load-Error Injection jointly improve global and local MoE load balance, cutting MaxVio up to 19.4% and 25% respectively at comparable model quality.

    Editor's note

    Must-read. For the first time advances the global quantiles of Quantile Balancing from rank averaging/histogram approximation to exact computation: two-pass BF16 radix selection recovers global batch empirical quantiles, with communication cost independent of token count and invariant to data partitioning; also proposes Load-Error Injection, using the constant diagonal Jacobian of unnormalized router scores as an STE proxy to inject local load errors, avoiding GShard's cross-expert coupling. Compared to Issue 1's Kimi K3 histogram approximation and Issue 5's Router Prior Bias soft anchoring, this paper provides provable increments in global quantile exactness and reports global and local MaxVio with dense baselines on 7.5B/500B tokens.

  3. Pure RL on executed outcomes trains a 122B-parameter MoE terminal agent to 64.0% on Terminal-Bench 2.1, surpassing GPT-5.4 with only 10B active parameters.

    Editor's note

    Must-read. For the first time decomposes training-inference consistency into two orthogonal axes—token fidelity (TITO) and routing fidelity (R3)—in agentic RL at 122B scale MoE: TITO uses token-in/token-out concatenation to ensure per-token alignment, R3 records and replays actual expert selections per layer during rollout (not just logits). Compared to Issue 1's PR² predictive replay, this paper directly records actual sampled routing, reducing the log-prob gap from 0.021 to 0.013 on long-horizon tasks with 300+ tool calls, marking the first validation of routing replay at frontier scale and in agentic scenarios.

  4. MegaFlux dynamically replicates experts across GPUs, pipelining weight transfers and gradient reductions to achieve 1.45x forward and 1.28x backward speedups over fixed placement.

    Editor's note

    Must-read. For the first time directly pipelines the weight transfer and gradient reduction caused by dynamic expert replication into the persistent scheduling of MoE megakernels: using phase-level weight-ready gating for forward, gradient tile priority production, and reduction-overlapped backward, rather than executing replication in stages outside the megakernel as in Issue 1's UltraEP. Compared to Issue 6's PipelinedLLEP memory-side mitigation, this paper pipelines replication costs from the execution side, reporting full routing statistics and dense baselines across 147 configurations on B200.

  5. Routing drift in merged MoE LLMs is mostly representation-induced, not router-parameter-induced, and fails to predict task-level recoverable loss, invalidating source-route repair as beneficial.

    Editor's note

    Worth reading. For the first time systematically falsifies "routing drift implies routing failure" in MoE model merging scenarios: through cross-interventions, attributes most expert reassignments to representation shifts rather than router parameter changes, and proves that structural distance (e.g., JS divergence) cannot predict source routing recovery gains (AUROC 0.47–0.52). Compared to Issue 5's Router Prior Bias and Issue 2's PR² training-inference consistency, this paper extends whether routing drift constitutes failure to merging scenarios, proposing intervention-relative recoverable loss as a task-level diagnostic standard, setting a new burden of proof for routing repair methods.

  6. Cobalt exploits expert co-activation to co-locate frequently paired experts, cutting cross-node traffic up to 99% and achieving up to 2.41x speedup in distributed MoE training.

    Editor's note

    Worth reading. For the first time uses "expert co-activation" as an explicit signal for cross-node communication reduction: under node-level token deduplication, co-activated expert pairs can share the same cross-node transmission, using Ochiai similarity to advance layout planning from single-expert statistics to expert-pair statistics. Compared to Issue 4's Director online lookahead placement, the increment here lies in the expert-pair co-activation signal itself, reporting end-to-end speedup and cross-node traffic reduction with dense baselines on 32×B200.