Summary (Overview)
- This paper studies the underexplored co-design of agentic reinforcement learning (RL) and mixture-of-experts (MoE) expert selection, showing that standard RL algorithms ignore the natural routing structure in agentic trajectories, limiting task performance and inference efficiency.
- The authors identify operation-grouped expert specialization (turns with semantically similar operations like READ, CREATE, UPDATE share expert distributions) and temporal routing consistency (adjacent tokens within a field reuse experts) as key structural patterns in off-the-shelf MoE models.
- They propose a hierarchical routing control framework with: (1) turn-level mutual information (MI) maximization between router distributions and operation labels, (2) token-level selective local consistency regularization, and (3) an entropy-gated control mechanism to stabilize training.
- The framework achieves over 10-point improvements in success rate across all evaluated benchmarks (AppWorld and AutomationBench), improves inference throughput by up to +41.1%, and stabilizes long-term MoE RL training.
- The method requires no architectural modifications, preserves the native top-k router, and is compatible with multiple RL algorithms (PPO, GRPO, LOOP, GiGPO).
Introduction and Theoretical Foundation
Background and Motivation
Long-horizon LLM agents are increasingly built using sparse mixture-of-experts (MoE) architectures, which activate only a fraction of parameters per token. While external scaffolding (e.g., ReAct) enables complex agentic behaviors, and MoE architectures enable efficient scaling, the co-design of agentic frameworks and MoE routing remains underexplored.
Agentic trajectories naturally alternate between:
- Thinking fields (): textual reasoning
- Tool-use fields (): executable code or tool calls
- Environment feedback (): returned observations
Each turn performs an identifiable semantic operation (e.g., READ, CREATE, UPDATE). The authors observe that MoE routers already exhibit operation-grouped specialization: turns with similar operations share expert distributions, and adjacent tokens within a field reuse experts. However, standard RL ignores this structure, injecting noise into routing statistics during training.
Key Theoretical Insight
The authors hypothesize that aligning expert routing with the semantic structure of trajectories—maintaining consistency within fields and specializing experts for distinct operations—allows the model to utilize expert capacity more effectively.
Methodology
1. Turn-Level Operation-Aware Control
The turn-level control maximizes mutual information (MI) between the turn-level expert selection distribution and operation labels. For turn , the turn-level expert distribution is:
where is the token indices of turn and is the router probability for expert at token . The MI between expert identity and operation presence is:
The stabilized turn-level MI objective is:
with and . The MI decomposition shows the objective complements load balancing: the marginal entropy term favors diverse overall usage, while the conditional entropy term favors concentrated usage within each operation state.
2. Token-Level Local Consistency Control
For adjacent tokens and within the same field, the expert-set gap is:
Given a threshold , the token-level loss is:
This is selective: it reinforces the previous token's expert set only when adjacent routes are already similar, preserving necessary routing changes at semantic and field boundaries. Pairs crossing field or turn boundaries are excluded.
3. Entropy-Gated Control
To prevent training collapse, the policy entropy is monitored:
The full objective becomes:
where . When the gate closes, routing-control gradients are detached while standard RL updates continue.
Empirical Validation / Results
Experimental Setup
- AppWorld: 90 training tasks, evaluated on test-normal and test-challenge splits
- AutomationBench: 480 public training tasks across six domains
- Models: Qwen3-30B-A3B-2507 (AppWorld), Qwen3.5-35B-A3B (AutomationBench)
- Hardware: 8 NVIDIA B200 GPUs
End-to-End Task Performance
Table 1: Agentic task performance (%). Bold and underline mark highest and second-highest scores per column.
| Method | AppWorld test-normal TGC | AppWorld test-normal SGC | AppWorld test-challenge TGC | AppWorld test-challenge SGC | AutomationBench |
|---|---|---|---|---|---|
| Base Model | 35.1 | 14.3 | 18.7 | 7.2 | 22.04 |
| PPO + Ours | 63.1 → 60.7 (2.4 ↓) | 37.5 → 42.9 (5.4 ↑) | 35.5 → 38.6 (3.1 ↑) | 16.5 → 20.9 (4.3 ↑) | 54.56 → 57.39 (2.83 ↑) |
| GRPO + Ours | 69.6 → 75.0 (5.4 ↑) | 46.4 → 55.4 (8.9 ↑) | 40.5 → 49.2 (8.6 ↑) | 22.3 → 34.5 (12.2 ↑) | 63.01 → 69.54 (6.53 ↑) |
| LOOP + Ours | 64.9 → 73.2 (8.3 ↑) | 46.4 → 58.9 (12.5 ↑) | 37.4 → 50.4 (12.9 ↑) | 21.6 → 30.9 (9.4 ↑) | 55.50 → 66.84 (11.34 ↑) |
| GiGPO + Ours | 52.4 → 57.7 (5.4 ↑) | 26.8 → 39.3 (12.5 ↑) | 32.9 → 37.4 (4.6 ↑) | 17.3 → 18.7 (1.4 ↑) | 70.01 → 70.24 (0.23 ↑) |
Key findings:
- LOOP + Ours achieves the largest improvements: +12.9 TGC on AppWorld challenge and +11.34 on AutomationBench
- GRPO + Ours shows consistent gains across all metrics, including +12.2 SGC on challenge
- PPO shows inconsistent improvement due to fewer trajectories per prompt, making MI estimation unstable
Routing Behavior Changes
- Token-level control increases within-field Jaccard similarity by 7.7%, indicating improved local consistency
- Turn-level control increases MI between expert distributions and operation labels by 23.5% in late-stage training
- Expert specialization (within-group vs. cross-group similarity gap) improves, confirming greater operation separability
Inference Efficiency
- Combined method increases throughput from 224 to 316 tokens/GPU/s (+41.1%)
- Wall time reduced from 224 to 174 seconds (22.3% reduction)
- Token-level control alone achieves the shortest wall time (133 seconds)
Entropy Gating Effectiveness
- Without entropy gating: policy entropy spikes and both evaluation splits collapse to zero around step 100
- With : collapse delayed to step 140
- With : training remains stable for all 200 steps
Theoretical and Practical Implications
Theoretical Implications
-
Agentic trajectory structure is a useful signal for MoE organization: The paper demonstrates that semantic operation labels provide effective supervision for expert specialization, complementing standard load-balancing losses.
-
MI objective vs. load balancing: The MI decomposition shows the proposed objective rewards operation-dependent usage, unlike standard load balancing which only encourages uniform utilization.
-
Stability-routing trade-off: Directly controlling MoE routers can destabilize RL training; policy entropy serves as an effective safeguard signal, connecting routing control to the broader literature on RL stability.
Practical Implications
-
No architectural changes required: The framework works with standard top-k MoE routers, making it easily deployable to existing models.
-
Algorithm-agnostic: Compatible with PPO, GRPO, LOOP, and GiGPO, suggesting broad applicability across RL paradigms.
-
Efficiency gains: Improved routing consistency translates to meaningful inference speedups (+41.1% throughput), reducing deployment costs for long-horizon agentic workloads.
Conclusion
This work demonstrates that agentic trajectory structure provides a useful signal for organizing MoE expert selection during post-training. The proposed hierarchical routing control framework combines:
- Operation-aware turn-level control (MI maximization) for expert specialization
- Selective token-level consistency for local routing stability
- Entropy-gated control for training stability
The framework consistently improves task performance across multiple RL algorithms and benchmarks, induces the intended routing structure, stabilizes long-term training, and improves inference efficiency.
Future Directions
- Extending routing control to other MoE architectures and scales
- Exploring richer operation taxonomies beyond the current grouping rules
- Investigating the interaction between routing control and other auxiliary objectives (e.g., load balancing) more deeply
- Applying the framework to other agentic domains beyond AppWorld and AutomationBench
Related papers
- Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild
Agent harnesses remain substantially undertested, with less than half of LLM-dependent code covered, and HarnessTester's contract-aware test generation boosts coverage and mutation scores by up to 95% while finding 88 previously-unknown bugs.
- Maintaining Benchmarks Against Increasingly Capable Agents: Detection and Remediation of Unearned Passes
A process-verification framework detects and patches reward hacking in agentic benchmarks, showing violation rates rise then fall across model generations and that replay tests alone cannot confirm repair.
- One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents
Co-installed coding-agent skills that do the same job reduce the installed skill's usage by 19.9 percentage points without lowering task completion, a conflict decided at the first skill read.