Summary (Overview)

  • This paper studies the underexplored co-design of agentic reinforcement learning (RL) and mixture-of-experts (MoE) expert selection, showing that standard RL algorithms ignore the natural routing structure in agentic trajectories, limiting task performance and inference efficiency.
  • The authors identify operation-grouped expert specialization (turns with semantically similar operations like READ, CREATE, UPDATE share expert distributions) and temporal routing consistency (adjacent tokens within a field reuse experts) as key structural patterns in off-the-shelf MoE models.
  • They propose a hierarchical routing control framework with: (1) turn-level mutual information (MI) maximization between router distributions and operation labels, (2) token-level selective local consistency regularization, and (3) an entropy-gated control mechanism to stabilize training.
  • The framework achieves over 10-point improvements in success rate across all evaluated benchmarks (AppWorld and AutomationBench), improves inference throughput by up to +41.1%, and stabilizes long-term MoE RL training.
  • The method requires no architectural modifications, preserves the native top-k router, and is compatible with multiple RL algorithms (PPO, GRPO, LOOP, GiGPO).

Introduction and Theoretical Foundation

Background and Motivation

Long-horizon LLM agents are increasingly built using sparse mixture-of-experts (MoE) architectures, which activate only a fraction of parameters per token. While external scaffolding (e.g., ReAct) enables complex agentic behaviors, and MoE architectures enable efficient scaling, the co-design of agentic frameworks and MoE routing remains underexplored.

Agentic trajectories naturally alternate between:

  • Thinking fields (yithinky_i^{\text{think}}): textual reasoning
  • Tool-use fields (yitooly_i^{\text{tool}}): executable code or tool calls
  • Environment feedback (viv_i): returned observations

Each turn performs an identifiable semantic operation (e.g., READ, CREATE, UPDATE). The authors observe that MoE routers already exhibit operation-grouped specialization: turns with similar operations share expert distributions, and adjacent tokens within a field reuse experts. However, standard RL ignores this structure, injecting noise into routing statistics during training.

Key Theoretical Insight

The authors hypothesize that aligning expert routing with the semantic structure of trajectories—maintaining consistency within fields and specializing experts for distinct operations—allows the model to utilize expert capacity more effectively.


Methodology

1. Turn-Level Operation-Aware Control

The turn-level control maximizes mutual information (MI) between the turn-level expert selection distribution and operation labels. For turn ii, the turn-level expert distribution is:

qi(e)=1∣Ti∣∑j∈Tipj(e)\mathbf{q}_i(e) = \frac{1}{|\mathcal{T}_i|} \sum_{j \in \mathcal{T}_i} \mathbf{p}_j(e)

where Ti\mathcal{T}_i is the token indices of turn ii and pj(e)\mathbf{p}_j(e) is the router probability for expert ee at token jj. The MI between expert identity EE and operation presence YoY_o is:

I(E;Yo)=KL(Po(E,Yo)∥P(E)⊗Po(Yo))(3.1)\mathcal{I}(E; Y_o) = \text{KL}\left(P_o(E, Y_o) \| P(E) \otimes P_o(Y_o)\right) \tag{3.1}

The stabilized turn-level MI objective is:

Jturn=1∣O∣∑o∈O∑e∑y∈{0,1}P^o(e,y)⋅clip(log⁡P^o(e,y)+ϵP^(e)P^o(y)+ϵ,−hc,hc)(3.3)J_{\text{turn}} = \frac{1}{|\mathcal{O}|} \sum_{o \in \mathcal{O}} \sum_{e} \sum_{y \in \{0,1\}} \hat{P}_o(e, y) \cdot \text{clip}\left(\log \frac{\hat{P}_o(e, y) + \epsilon}{\hat{P}(e)\hat{P}_o(y) + \epsilon}, -h_c, h_c\right) \tag{3.3}

with ϵ=10−12\epsilon = 10^{-12} and hc=20h_c = 20. The MI decomposition I(E;Yo)=H(E)−H(E∣Yo)I(E; Y_o) = H(E) - H(E|Y_o) shows the objective complements load balancing: the marginal entropy term favors diverse overall usage, while the conditional entropy term favors concentrated usage within each operation state.

2. Token-Level Local Consistency Control

For adjacent tokens t−1t-1 and tt within the same field, the expert-set gap is:

δt=1−∣St−1∩St∣/∣St∣\delta_t = 1 - |\mathcal{S}_{t-1} \cap \mathcal{S}_t| / |\mathcal{S}_t|

Given a threshold hδ∈[0,1]h_\delta \in [0, 1], the token-level loss is:

ℓttoken=−I[δt<hδ]∣St−1∣∑e∈St−1log⁡pt[e](3.4)\ell_t^{\text{token}} = -\frac{\mathbb{I}[\delta_t < h_\delta]}{|\mathcal{S}_{t-1}|} \sum_{e \in \mathcal{S}_{t-1}} \log p_t[e] \tag{3.4}

This is selective: it reinforces the previous token's expert set only when adjacent routes are already similar, preserving necessary routing changes at semantic and field boundaries. Pairs crossing field or turn boundaries are excluded.

3. Entropy-Gated Control

To prevent training collapse, the policy entropy is monitored:

Ht=−∑v∈Vπθ(v∣τ<t)⋅log⁡πθ(v∣τ<t)(3.5)\mathcal{H}_t = -\sum_{v \in \mathbb{V}} \pi_{\boldsymbol{\theta}}(v | \tau_{<t}) \cdot \log \pi_{\boldsymbol{\theta}}(v | \tau_{<t}) \tag{3.5}

The full objective becomes:

J=JRL+g⋅(λturnJturn+λtokenJtoken)(3.6)J = J_{\text{RL}} + g \cdot (\lambda_{\text{turn}} J_{\text{turn}} + \lambda_{\text{token}} J_{\text{token}}) \tag{3.6}

where g=I[Hlow≤H≤Hhigh]g = \mathbb{I}[H_{\text{low}} \le \mathcal{H} \le H_{\text{high}}]. When the gate closes, routing-control gradients are detached while standard RL updates continue.


Empirical Validation / Results

Experimental Setup

  • AppWorld: 90 training tasks, evaluated on test-normal and test-challenge splits
  • AutomationBench: 480 public training tasks across six domains
  • Models: Qwen3-30B-A3B-2507 (AppWorld), Qwen3.5-35B-A3B (AutomationBench)
  • Hardware: 8 NVIDIA B200 GPUs

End-to-End Task Performance

Table 1: Agentic task performance (%). Bold and underline mark highest and second-highest scores per column.

MethodAppWorld test-normal TGCAppWorld test-normal SGCAppWorld test-challenge TGCAppWorld test-challenge SGCAutomationBench
Base Model35.114.318.77.222.04
PPO + Ours63.1 → 60.7 (2.4 ↓)37.5 → 42.9 (5.4 ↑)35.5 → 38.6 (3.1 ↑)16.5 → 20.9 (4.3 ↑)54.56 → 57.39 (2.83 ↑)
GRPO + Ours69.6 → 75.0 (5.4 ↑)46.4 → 55.4 (8.9 ↑)40.5 → 49.2 (8.6 ↑)22.3 → 34.5 (12.2 ↑)63.01 → 69.54 (6.53 ↑)
LOOP + Ours64.9 → 73.2 (8.3 ↑)46.4 → 58.9 (12.5 ↑)37.4 → 50.4 (12.9 ↑)21.6 → 30.9 (9.4 ↑)55.50 → 66.84 (11.34 ↑)
GiGPO + Ours52.4 → 57.7 (5.4 ↑)26.8 → 39.3 (12.5 ↑)32.9 → 37.4 (4.6 ↑)17.3 → 18.7 (1.4 ↑)70.01 → 70.24 (0.23 ↑)

Key findings:

  • LOOP + Ours achieves the largest improvements: +12.9 TGC on AppWorld challenge and +11.34 on AutomationBench
  • GRPO + Ours shows consistent gains across all metrics, including +12.2 SGC on challenge
  • PPO shows inconsistent improvement due to fewer trajectories per prompt, making MI estimation unstable

Routing Behavior Changes

  • Token-level control increases within-field Jaccard similarity by 7.7%, indicating improved local consistency
  • Turn-level control increases MI between expert distributions and operation labels by 23.5% in late-stage training
  • Expert specialization (within-group vs. cross-group similarity gap) improves, confirming greater operation separability

Inference Efficiency

  • Combined method increases throughput from 224 to 316 tokens/GPU/s (+41.1%)
  • Wall time reduced from 224 to 174 seconds (22.3% reduction)
  • Token-level control alone achieves the shortest wall time (133 seconds)

Entropy Gating Effectiveness

  • Without entropy gating: policy entropy spikes and both evaluation splits collapse to zero around step 100
  • With Hhigh=0.4H_{\text{high}} = 0.4: collapse delayed to step 140
  • With Hhigh=0.2H_{\text{high}} = 0.2: training remains stable for all 200 steps

Theoretical and Practical Implications

Theoretical Implications

  1. Agentic trajectory structure is a useful signal for MoE organization: The paper demonstrates that semantic operation labels provide effective supervision for expert specialization, complementing standard load-balancing losses.

  2. MI objective vs. load balancing: The MI decomposition I(E;Yo)=H(E)−H(E∣Yo)I(E; Y_o) = H(E) - H(E|Y_o) shows the proposed objective rewards operation-dependent usage, unlike standard load balancing which only encourages uniform utilization.

  3. Stability-routing trade-off: Directly controlling MoE routers can destabilize RL training; policy entropy serves as an effective safeguard signal, connecting routing control to the broader literature on RL stability.

Practical Implications

  1. No architectural changes required: The framework works with standard top-k MoE routers, making it easily deployable to existing models.

  2. Algorithm-agnostic: Compatible with PPO, GRPO, LOOP, and GiGPO, suggesting broad applicability across RL paradigms.

  3. Efficiency gains: Improved routing consistency translates to meaningful inference speedups (+41.1% throughput), reducing deployment costs for long-horizon agentic workloads.


Conclusion

This work demonstrates that agentic trajectory structure provides a useful signal for organizing MoE expert selection during post-training. The proposed hierarchical routing control framework combines:

  1. Operation-aware turn-level control (MI maximization) for expert specialization
  2. Selective token-level consistency for local routing stability
  3. Entropy-gated control for training stability

The framework consistently improves task performance across multiple RL algorithms and benchmarks, induces the intended routing structure, stabilizes long-term training, and improves inference efficiency.

Future Directions

  • Extending routing control to other MoE architectures and scales
  • Exploring richer operation taxonomies beyond the current grouping rules
  • Investigating the interaction between routing control and other auxiliary objectives (e.g., load balancing) more deeply
  • Applying the framework to other agentic domains beyond AppWorld and AutomationBench

Related papers