Summary (Overview)

  • RelayMoE is a ring-based Mixture-of-Experts (MoE) execution model that eliminates full top-k-expanded dispatch buffers by circulating expert weights or token activations through a ring of GPUs, computing partial results locally at each hop.
  • The approach reduces peak GPU memory during MoE execution by up to 7.6× on average at EP=64 (reaching 13.4× for top-16 routing), addressing the memory bottleneck that limits training of increasingly large MoE models.
  • RelayMoE achieves a 2× average speedup over Megatron-LM's all-to-all dispatcher in single-layer MoE experiments at EP=8, and improves full-model training throughput by up to 2.02× under the same GPU memory budget.
  • The design extends the largest tested trainable sequence length by up to 2.85× and supports larger micro-batches (18–32 vs. 12 for baselines) by converting memory savings into training capacity.
  • The system features adaptive routing (expert vs. token circulation) based on communication volume, memory-efficient MoE recomputation during backward, and selective attention recomputation to further improve throughput.

Introduction and Theoretical Foundation

Background

Mixture-of-Experts (MoE) architectures have become dominant in large language model training by activating only a subset of experts per token, achieving stronger quality at a fraction of dense model compute cost. However, frontier MoE models are scaling along multiple dimensions simultaneously—more experts, higher top-k routing, longer sequences, and wider expert FFN layers—each independently increasing GPU memory requirements.

The Memory Bottleneck

The paper identifies a critical insight: peak memory in distributed MoE training is dominated by the MoE block, not attention. The standard all-to-all dispatcher requires constructing the full top-k-expanded buffer before communication can begin, meaning every intermediate buffer in the dispatch pipeline is scaled by the top-k factor:

k×(base activation tensor)k \times \text{(base activation tensor)}

Table 1 shows how this compounds with architectural trends:

Model# ExpertsTop-kMax ContextBuffer Multiplier
Mixtral 8x22B8232K2×
DeepSeek-V32568128K8×
Qwen3-30B-A3B1288131K8×
Kimi K23848128K8×
Qwen3-Next-80B-A3B51210128K10×
Qwen3.5-35B-A3B2568262K8×

Combined with growing sequence lengths (32K→262K), peak memory from MoE dispatch has grown by an order of magnitude. While attention memory has been effectively reduced through memory-efficient algorithms (FlashAttention, hybrid linear-attention designs), MoE dispatch remains the dominant contributor to peak memory, creating transient memory spikes that determine the out-of-memory (OOM) boundary.

Key Observation

The paper profiles a Qwen3-style model and shows that MoE blocks create memory spikes up to 51 GB while memory between spikes remains around 14 GB. The top-k expansion multiplies tokens per GPU by 8×, producing dispatch buffers that cause the spikes. Two factors drive peak memory:

  1. Each intermediate activation in the dispatch pipeline is proportional to tokens per GPU × top-k × hidden dimension
  2. Multiple such activations coexist in memory during dispatch pipeline execution

Methodology

Ring-Based MoE Dispatch

RelayMoE circulates data through the expert-parallel group in a ring instead of gathering all routed tokens in a single collective step. The core principle: compute locally as data circulates.

Expert Routing (Forward)

At each hop, a GPU:

  1. Uses the routing map to find local tokens that selected the visiting expert partition
  2. Locally permutes relevant tokens and executes expert GEMM
  3. Weights outputs by router probabilities and accumulates into running output

After NN hops across NN GPUs, every expert partition has visited every GPU, completing the output activation.

Communication-Computation Overlap

  • Forward: While computing with the current expert partition, asynchronously prefetch the next partition's weights via P2P communication
  • Backward: Overlap local recomputation and gradient computation with merged transfers of expert weights and accumulated weight gradients

The per-hop time approaches max⁡(computation,transfer)\max(\text{computation}, \text{transfer}) rather than their sum.

Adaptive Routing: Experts vs. Tokens

RelayMoE selects between two modes based on communication volume:

Routing ExpertRouting Token
ResidentTokens: B×S×HB \times S \times HExperts: Elocal×WE_{local} \times W
Comm. (fwd)Elocal×WE_{local} \times WB×S×(H+M)B \times S \times (H + M)
Comm. (bwd)2×Elocal×W2 \times E_{local} \times W2×B×S×(H+M)2 \times B \times S \times (H + M)

Where BB = micro-batch, SS = sequence length, HH = hidden dimension, WW = single expert weight size, MM = routing metadata size per token.

Decision rule: Expert routing is preferred when token activations are large relative to expert weights; token routing is preferred when expert weights are large relative to token activations. The choice depends on the full training configuration (including SP, CP, and ETP settings), not just model architecture.

Memory-Efficient MoE Recomputation

The ring structure naturally supports recomputation within each backward hop:

  • Reconstruct expert intermediates using saved local tokens + visiting expert partition
  • Compute and accumulate gradients immediately
  • Release reconstructed intermediates before the next hop

This prevents expert intermediates from accumulating across hops, with peak memory determined by the largest hop rather than the sum.

Selective Attention Recomputation

The saved memory is used to selectively retain attention activations. For each attention layer ii and recomputation option jj:

  • Memory cost: mi,jm_{i,j} (retained activations)
  • Compute saving: ci,jc_{i,j} (avoided backward recomputation)

RelayMoE ranks options by compute saving per unit memory ci,j/mi,jc_{i,j}/m_{i,j} and greedily selects within the memory budget.

Empirical Validation / Results

Experimental Setup

  • Platform: MareNostrum 5, NVIDIA H100 GPUs (64 GB HBM2e), NVLink 4.0 intra-node, NDR InfiniBand inter-node
  • Implementation: 2.3K lines of code on top of Megatron-LM
  • Baselines: Megatron-LM all-to-all, FasterMoE, Tutel, DeepEP

MoE Layer Performance

  • 2× average speedup over all-to-all at EP=8
  • Speedup increases with sequence length: 1.77× at 40K → 2.4× at 98K for E128
  • 4.03× faster than FasterMoE on average
  • 2.94× faster than Tutel on average
  • At EP=64: speedups up to 27× over FasterMoE and 16× over Tutel
  • Peak memory reduction: 3.8× at EP=8, 7.6× at EP=64 (13.4× for top-16)

Full-Model Training

  • Throughput improvements on production models (Qwen3-30B-A3B, Qwen3.5-35B-A3B, Qwen2-57B-A14B):
    • Up to 2.02× speedup (K16 at 40K)
    • Largest tested sequence length extended from 393K to 786K (2× improvement)
    • Qwen2-57B trains at 384K where all-to-all runs OOM

Memory Savings Utilization

Use of Saved MemoryResult
Reduced recomputationUp to 2.02× throughput improvement
Larger micro-batches18–32 vs. 12 for all-to-all
Longer sequencesUp to 2.85× extension (90K–150K vs. 65K max for all-to-all)

Design Analysis

  • Communication overlap: Reduces latency from 295.9 ms → 203.4 ms (naive → overlapped)
  • GroupedGEMM: Further reduces to 175.8 ms
  • Token routing: Outperforms expert routing by 2.1–3.7× on wide-expert configurations (384 experts), and runs at sequence lengths where all-to-all encounters OOM
  • Training correctness: Loss curves indistinguishable from all-to-all over 1500 iterations

Theoretical and Practical Implications

Theoretical Contributions

  1. Memory-computation tradeoff reframed: The paper demonstrates that the all-to-all dispatcher's memory cost is not inherent to MoE computation but an artifact of execution order. By reorganizing computation as a ring, the same mathematical operations (token–expert GEMM pairs) can be computed with dramatically lower peak memory.

  2. Mathematical equivalence: RelayMoE is provably equivalent to all-to-all dispatch—each token–expert pair computes the same GEMM with the same weights and router probabilities; only the execution order differs.

  3. Routing mode selection: Provides a principled framework for choosing between expert and token circulation based on communication volume, with clear formulas for when each mode is preferable.

Practical Implications

  1. Extended training capacity: Enables training longer sequences and larger batches on the same hardware, directly addressing the OOM boundary that limits MoE training.

  2. Complementary to existing optimizations: Works alongside attention memory optimizations (FlashAttention, sequence/context parallelism) and can be combined with optimized dispatch kernels.

  3. Production readiness: Demonstrated on production models (Qwen2-57B, Qwen3-30B, Qwen3.5-35B) with state-of-the-art parallelization strategies (MoE Parallel Folding).

  4. Dropless execution: Eliminates the need for token dropping and capacity-factor padding, simplifying training pipelines.

Conclusion

RelayMoE presents a fundamental redesign of MoE execution in distributed training, replacing the memory-hungry all-to-all dispatch pipeline with a ring-based model that computes locally as data circulates. The key insights are:

  1. Memory bottleneck identification: MoE blocks, not attention, determine the OOM boundary in modern MoE training
  2. Ring-based execution: Avoids full top-k-expanded buffers while maintaining mathematical equivalence
  3. Memory-to-throughput conversion: Saved memory can be strategically deployed (larger batches, longer sequences, or reduced recomputation) for maximum training benefit

The work opens several future directions:

  • Hierarchical ring designs to better utilize intra-node NVLink bandwidth vs. inter-node InfiniBand
  • Integration with other memory reduction techniques (offloading, model-state partitioning)
  • Extension to inference workloads where similar memory constraints apply
  • Adaptive hop schedules that respond to dynamic routing patterns

The results—2× average speedup over all-to-all, 7.6× memory reduction, and 2.85× sequence length extension—demonstrate that rethinking the fundamental execution model of MoE training can yield substantial practical gains beyond what communication optimization alone can achieve.

Related papers