# Structuring MoE Expert Selection for Agentic Reinforcement Learning

> Hierarchical routing control that aligns MoE expert selection with agentic trajectory structure improves success rates by over 10 points and boosts inference throughput up to 41.1% without architectural changes.

- **Source:** [arXiv](https://arxiv.org/abs/2610.07332)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/mBnVl3
- **Whiteboard:** https://picx.dev/p/mBnVl3/image

## Summary

## Summary (Overview)

- This paper studies the underexplored co-design of agentic reinforcement learning (RL) and mixture-of-experts (MoE) expert selection, showing that standard RL algorithms ignore the natural routing structure in agentic trajectories, limiting task performance and inference efficiency.
- The authors identify **operation-grouped expert specialization** (turns with semantically similar operations like READ, CREATE, UPDATE share expert distributions) and **temporal routing consistency** (adjacent tokens within a field reuse experts) as key structural patterns in off-the-shelf MoE models.
- They propose a **hierarchical routing control framework** with: (1) turn-level mutual information (MI) maximization between router distributions and operation labels, (2) token-level selective local consistency regularization, and (3) an entropy-gated control mechanism to stabilize training.
- The framework achieves **over 10-point improvements in success rate** across all evaluated benchmarks (AppWorld and AutomationBench), improves inference throughput by up to +41.1%, and stabilizes long-term MoE RL training.
- The method requires **no architectural modifications**, preserves the native top-k router, and is compatible with multiple RL algorithms (PPO, GRPO, LOOP, GiGPO).

---

## Introduction and Theoretical Foundation

### Background and Motivation

Long-horizon LLM agents are increasingly built using sparse mixture-of-experts (MoE) architectures, which activate only a fraction of parameters per token. While external scaffolding (e.g., ReAct) enables complex agentic behaviors, and MoE architectures enable efficient scaling, **the co-design of agentic frameworks and MoE routing remains underexplored**.

Agentic trajectories naturally alternate between:
- **Thinking fields** ($y_i^{\text{think}}$): textual reasoning
- **Tool-use fields** ($y_i^{\text{tool}}$): executable code or tool calls
- **Environment feedback** ($v_i$): returned observations

Each turn performs an identifiable semantic operation (e.g., READ, CREATE, UPDATE). The authors observe that MoE routers already exhibit **operation-grouped specialization**: turns with similar operations share expert distributions, and adjacent tokens within a field reuse experts. However, standard RL ignores this structure, injecting noise into routing statistics during training.

### Key Theoretical Insight

The authors hypothesize that **aligning expert routing with the semantic structure of trajectories**—maintaining consistency within fields and specializing experts for distinct operations—allows the model to utilize expert capacity more effectively.

---

## Methodology

### 1. Turn-Level Operation-Aware Control

The turn-level control maximizes **mutual information (MI)** between the turn-level expert selection distribution and operation labels. For turn $i$, the turn-level expert distribution is:

$$\mathbf{q}_i(e) = \frac{1}{|\mathcal{T}_i|} \sum_{j \in \mathcal{T}_i} \mathbf{p}_j(e)$$

where $\mathcal{T}_i$ is the token indices of turn $i$ and $\mathbf{p}_j(e)$ is the router probability for expert $e$ at token $j$. The MI between expert identity $E$ and operation presence $Y_o$ is:

$$\mathcal{I}(E; Y_o) = \text{KL}\left(P_o(E, Y_o) \| P(E) \otimes P_o(Y_o)\right) \tag{3.1}$$

The stabilized turn-level MI objective is:

$$J_{\text{turn}} = \frac{1}{|\mathcal{O}|} \sum_{o \in \mathcal{O}} \sum_{e} \sum_{y \in \{0,1\}} \hat{P}_o(e, y) \cdot \text{clip}\left(\log \frac{\hat{P}_o(e, y) + \epsilon}{\hat{P}(e)\hat{P}_o(y) + \epsilon}, -h_c, h_c\right) \tag{3.3}$$

with $\epsilon = 10^{-12}$ and $h_c = 20$. The MI decomposition $I(E; Y_o) = H(E) - H(E|Y_o)$ shows the objective complements load balancing: the marginal entropy term favors diverse overall usage, while the conditional entropy term favors concentrated usage within each operation state.

### 2. Token-Level Local Consistency Control

For adjacent tokens $t-1$ and $t$ within the same field, the expert-set gap is:

$$\delta_t = 1 - |\mathcal{S}_{t-1} \cap \mathcal{S}_t| / |\mathcal{S}_t|$$

Given a threshold $h_\delta \in [0, 1]$, the token-level loss is:

$$\ell_t^{\text{token}} = -\frac{\mathbb{I}[\delta_t < h_\delta]}{|\mathcal{S}_{t-1}|} \sum_{e \in \mathcal{S}_{t-1}} \log p_t[e] \tag{3.4}$$

This is **selective**: it reinforces the previous token's expert set only when adjacent routes are already similar, preserving necessary routing changes at semantic and field boundaries. Pairs crossing field or turn boundaries are excluded.

### 3. Entropy-Gated Control

To prevent training collapse, the policy entropy is monitored:

$$\mathcal{H}_t = -\sum_{v \in \mathbb{V}} \pi_{\boldsymbol{\theta}}(v | \tau_{<t}) \cdot \log \pi_{\boldsymbol{\theta}}(v | \tau_{<t}) \tag{3.5}$$

The full objective becomes:

$$J = J_{\text{RL}} + g \cdot (\lambda_{\text{turn}} J_{\text{turn}} + \lambda_{\text{token}} J_{\text{token}}) \tag{3.6}$$

where $g = \mathbb{I}[H_{\text{low}} \le \mathcal{H} \le H_{\text{high}}]$. When the gate closes, routing-control gradients are detached while standard RL updates continue.

---

## Empirical Validation / Results

### Experimental Setup

- **AppWorld**: 90 training tasks, evaluated on test-normal and test-challenge splits
- **AutomationBench**: 480 public training tasks across six domains
- **Models**: Qwen3-30B-A3B-2507 (AppWorld), Qwen3.5-35B-A3B (AutomationBench)
- **Hardware**: 8 NVIDIA B200 GPUs

### End-to-End Task Performance

**Table 1: Agentic task performance (%).** Bold and underline mark highest and second-highest scores per column.

| Method | AppWorld test-normal TGC | AppWorld test-normal SGC | AppWorld test-challenge TGC | AppWorld test-challenge SGC | AutomationBench |
|---|---|---|---|---|---|
| Base Model | 35.1 | 14.3 | 18.7 | 7.2 | 22.04 |
| PPO + Ours | 63.1 → 60.7 (2.4 ↓) | 37.5 → 42.9 (5.4 ↑) | 35.5 → 38.6 (3.1 ↑) | 16.5 → 20.9 (4.3 ↑) | 54.56 → 57.39 (2.83 ↑) |
| GRPO + Ours | 69.6 → 75.0 (5.4 ↑) | 46.4 → 55.4 (8.9 ↑) | 40.5 → 49.2 (8.6 ↑) | 22.3 → 34.5 (12.2 ↑) | 63.01 → 69.54 (6.53 ↑) |
| LOOP + Ours | 64.9 → 73.2 (8.3 ↑) | 46.4 → 58.9 (12.5 ↑) | 37.4 → 50.4 (12.9 ↑) | 21.6 → 30.9 (9.4 ↑) | 55.50 → 66.84 (11.34 ↑) |
| GiGPO + Ours | 52.4 → 57.7 (5.4 ↑) | 26.8 → 39.3 (12.5 ↑) | 32.9 → 37.4 (4.6 ↑) | 17.3 → 18.7 (1.4 ↑) | 70.01 → 70.24 (0.23 ↑) |

Key findings:
- **LOOP + Ours** achieves the largest improvements: +12.9 TGC on AppWorld challenge and +11.34 on AutomationBench
- **GRPO + Ours** shows consistent gains across all metrics, including +12.2 SGC on challenge
- PPO shows inconsistent improvement due to fewer trajectories per prompt, making MI estimation unstable

### Routing Behavior Changes

- **Token-level control** increases within-field Jaccard similarity by 7.7%, indicating improved local consistency
- **Turn-level control** increases MI between expert distributions and operation labels by 23.5% in late-stage training
- **Expert specialization** (within-group vs. cross-group similarity gap) improves, confirming greater operation separability

### Inference Efficiency

- Combined method increases throughput from **224 to 316 tokens/GPU/s (+41.1%)**
- Wall time reduced from 224 to 174 seconds (22.3% reduction)
- Token-level control alone achieves the shortest wall time (133 seconds)

### Entropy Gating Effectiveness

- Without entropy gating: policy entropy spikes and both evaluation splits collapse to zero around step 100
- With $H_{\text{high}} = 0.4$: collapse delayed to step 140
- With $H_{\text{high}} = 0.2$: training remains stable for all 200 steps

---

## Theoretical and Practical Implications

### Theoretical Implications

1. **Agentic trajectory structure is a useful signal for MoE organization**: The paper demonstrates that semantic operation labels provide effective supervision for expert specialization, complementing standard load-balancing losses.

2. **MI objective vs. load balancing**: The MI decomposition $I(E; Y_o) = H(E) - H(E|Y_o)$ shows the proposed objective rewards operation-dependent usage, unlike standard load balancing which only encourages uniform utilization.

3. **Stability-routing trade-off**: Directly controlling MoE routers can destabilize RL training; policy entropy serves as an effective safeguard signal, connecting routing control to the broader literature on RL stability.

### Practical Implications

1. **No architectural changes required**: The framework works with standard top-k MoE routers, making it easily deployable to existing models.

2. **Algorithm-agnostic**: Compatible with PPO, GRPO, LOOP, and GiGPO, suggesting broad applicability across RL paradigms.

3. **Efficiency gains**: Improved routing consistency translates to meaningful inference speedups (+41.1% throughput), reducing deployment costs for long-horizon agentic workloads.

---

## Conclusion

This work demonstrates that **agentic trajectory structure provides a useful signal for organizing MoE expert selection during post-training**. The proposed hierarchical routing control framework combines:

1. **Operation-aware turn-level control** (MI maximization) for expert specialization
2. **Selective token-level consistency** for local routing stability
3. **Entropy-gated control** for training stability

The framework consistently improves task performance across multiple RL algorithms and benchmarks, induces the intended routing structure, stabilizes long-term training, and improves inference efficiency.

### Future Directions

- Extending routing control to other MoE architectures and scales
- Exploring richer operation taxonomies beyond the current grouping rules
- Investigating the interaction between routing control and other auxiliary objectives (e.g., load balancing) more deeply
- Applying the framework to other agentic domains beyond AppWorld and AutomationBench

---

_Markdown view of https://picx.dev/p/mBnVl3, served by PicX — AI-generated visual whiteboard summaries of research papers._
