Full text not available for this paper
Summary (Overview)
- MegaFlux is a system that enables dynamic expert replication within persistent Mixture-of-Experts (MoE) megakernels to address GPU stragglers caused by routing skew, without altering router outputs or token assignments.
- An on-device planner jointly selects replica locations and assigns tile-aligned token blocks under a per-GPU replica budget, using bounded greedy search for load balancing.
- The forward megakernel pipelines replica-weight transfers with expert computation by exploiting phase-level weight readiness (FC1 can start before FC2 weights arrive).
- The backward megakernel pipelines data-gradient computation, weight-gradient reduction, requantization, and token movement, prioritizing gradient production for replicated experts to enable early reduction.
- On eight NVIDIA B200 GPUs, MegaFlux achieves geometric-mean speedups of 1.45× (forward) and 1.28× (backward) over fixed placement, peaking at 2.14× and 2.64×, with end-to-end prefill speedups of 1.13–1.26× when integrated into vLLM.
Introduction and Theoretical Foundation
Background and Motivation
Mixture-of-Experts (MoE) has become a common scaling strategy in large language models (DeepSeek-V3, Qwen3, GLM-4.5, Kimi K2). By activating only a few experts per token, MoE increases parameter capacity without proportionally increasing per-token computation. Expert parallelism (EP) partitions experts across GPUs, introducing two system challenges:
- Communication for moving routed activations
- Load imbalance when input-dependent routing concentrates work on experts hosted by only a few GPUs
Fixed-Placement Megakernels
Recent MoE megakernels fuse token dispatch, expert computation, and token return into persistent kernels, enabling fine-grained pipelining. However, expert ownership remains fixed: a GPU can execute an expert only if it holds that expert's weights. Megakernels make each GPU's assigned work run efficiently but do not change where that work can execute.
Routing Skew Problem
Input-dependent routing can concentrate work on a few experts, leaving some GPUs underutilized while hot-expert GPUs determine layer latency. The paper quantifies routing skew as:
where is the token count routed to expert and is the total number of experts. Uniform routing gives ; larger values indicate more uneven loads.
In captured Qwen3 and OLMoE workloads, routing skew slows fixed-placement MegaMoE by 7.3–41.0%, motivating the need for dynamic expert replication.
Key Insight
Dynamic replication introduces new communication: replicas must receive expert weights to execute, and during training, their partial weight gradients must be reduced at expert owners. The challenge is realizing dynamic replication within the fine-grained communication-computation pipeline of MoE megakernels.
Methodology
1. Planner: Replica Placement and Token-Block Assignment
Problem Formulation
The planner partitions each expert's routed rows into token blocks of at most rows, matching the megakernel's tiled computation. The optimization problem is:
subject to:
where is the number of expert 's blocks assigned to GPU , is the block count, is the owner of expert , and is the per-GPU replica budget.
Bounded Greedy Search
Four greedy searches run in parallel, each starting from the no-replica assignment and shifting blocks from overloaded owners toward blocks per GPU. Searches combine:
- Donor priority: decreasing total overload or largest remaining home-expert load
- Fan-out: at most one new copy per expert per round, or multiple copies
For backward, a reduction-aware variant accounts for owner-side cost of replica-gradient reduction using the proxy:
where counts remote replicas of experts owned by GPU , and is a fixed penalty.
2. Executor: Pipelined Expert Replication
Forward
A replica can begin useful work before all weights arrive:
- FC1 requires only and an input token block
- FC2 requires and the completed activation block from FC1
Separate readiness for the two weight matrices overlaps FC1 with transfer. Communication warps progressively join the replica-weight queue after draining their token blocks.
Backward
The backward megakernel overlaps both replica materialization and gradient reduction with the existing DGrad-WGrad pipeline. Key mechanisms:
- Prioritized gradient production: WGrad preparation and gradient-tile production for replicated experts are prioritized across all participants
- Early gradient reduction: Owner-side workers pull and accumulate partials in FP32 in a fixed order as soon as they become available
- Overlapped requantization: Token-axis and requantization is scheduled in otherwise idle communication intervals
The expert computation follows:
where , , and denotes expert 's output for token .
Backward gradients:
Empirical Validation / Results
Performance Relative to Fixed Placement
| Metric | Forward | Backward |
|---|---|---|
| Geometric-mean speedup | 1.45× | 1.28× |
| Peak speedup (real-median) | 2.14× | 1.43× |
| Peak speedup (real-high) | 1.99× | 1.46× |
| Peak speedup (real-extreme) | 1.93× | 1.92× |
| Peak speedup (Zipf-κ3) | 1.53× | 1.75× |
| Peak speedup (Zipf-κ6) | 1.86× | 2.64× |
- Balanced routes remain near parity (0.985× forward, 0.999× backward geometric mean)
- Small skewed backward workloads can regress to 0.91×
- Gains strengthen as computation amortizes planning and replica-operation overhead
Decomposing Pipelined Expert Replication
Separate-stage replication reduces latency by up to 42.4% forward and 51.2% backward; pipelining provides an additional 13.2% and 26.7% reduction.
Hidden replica-operation cost (real-high routes):
| Operation | Hidden Fraction |
|---|---|
| Forward weight transfer | 56–76% |
| Backward transfer + reduction | 91–100% |
Comparison to Prior Art (BF16, E=256)
| Baseline | Forward Speedup | Backward Speedup |
|---|---|---|
| Megatron-Core | 2.31× | 1.53× |
| UltraEP | 1.74× | 2.35× |
| Mixture-of-Kittens (unmodified) | 1.84× | 1.37× |
| MoK (comparable-work estimate) | 1.70× | 1.26× |
Online Planner
- Planning takes 42.6–65.9 μs (E=128) and 49.8–72.0 μs (E=256)
- Planning-to-forward-latency ratio decreases from 9.17% to 0.14% as tokens/rank increase from 1K to 128K
- Planner is 1.31–2.35× faster than UltraEP's planner with CUDA-graph replay
End-to-End Evaluation (vLLM, DeepSeek-V4-Pro)
| Batch | 16K chunk | 32K chunk |
|---|---|---|
| 8 | 1.128× | 1.133× |
| 16 | 1.227× | 1.238× |
| 32 | 1.165× | 1.260× |
Theoretical and Practical Implications
Theoretical Contributions
-
Block-granular runtime expert replication: Formulates replication as a constrained optimization problem with tile-aligned token blocks, preserving token counts, expert selections, and routing weights.
-
Phase-level weight readiness: Demonstrates that replicas need not be fully materialized before computation begins—FC1 can overlap with transfer.
-
Pipelined gradient reduction: Shows that prioritizing gradient production for replicated experts across all participants enables early reduction that overlaps with remaining computation.
Practical Implications
- System design: MegaFlux demonstrates that persistent MoE execution can adapt to routing skew without being constrained by fixed expert placement, achieving significant speedups on real hardware.
- Workload amortization: Larger workloads amortize planning and replica-operation overhead, making replication most effective when substantial compute exists.
- Balanced workloads: Near-balanced assignments do not guarantee lower latency; exposed replica-operation costs may outweigh rebalancing benefits.
Limitations
- Evaluation is limited to a single NVLink domain; topology-aware, cross-node replication is left to future work
- A runtime gate to enable replication only when predicted savings exceed costs is suggested but not implemented
- The reduction-aware penalty is a heuristic; a more principled, workload-adaptive cost model is left to future work
Conclusion
MegaFlux resolves GPU stragglers due to work imbalance in persistent MoE execution via dynamic expert replication. An on-device planner assigns replicas and token blocks, while forward and backward megakernels pipeline replica-weight transfers and gradient reductions with expert computation. Across diverse routing skew and workload sizes on eight NVIDIA B200 GPUs, MegaFlux achieves geometric-mean speedups of 1.45× in forward and 1.28× in backward execution over fixed placement, peaking at 2.14× and 2.64×.
Future directions include:
- Topology-aware, cross-node replication
- Runtime gating to enable replication only when beneficial
- More principled cost models for reduction-aware planning
- Extending to additional workload types and model architectures
Related papers
- Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
A fixed LLM judge produces version-dependent errors, invalidating agent comparisons and transported calibration, so release decisions require paired audits, not judge-only scores.
- Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons
Fast learning-rate transfer holds at growing training horizons when T grows slower than sqrt(n), but requires nondegenerate first-order loss sensitivity to avoid spectral-dependent failures.
- Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
SuffixReplay enables fine-grained prefix caching in hybrid LLMs by reconstructing linear-attention states from sparse anchors, cutting storage by 2x and TTFT by up to 70%.