Hierarchical routing control lifts MoE agent RL success rates by over 10 points on AppWorld and AutomationBench
Related research and updatesSynopsis
This work systematically studies how agentic post-training interacts with MoE expert selection, observing that off-the-shelf MoE routers already show operation-grouped specialization (e.g., READ, UPDATE) that standard RL perturbs; the authors propose a hierarchical routing control framework that aligns turn-level expert selection with operation labels via mutual information, regularizes token-level adjacent consistency with a gap threshold, and adds entropy gating for stability, improving success rates by over 10 points across PPO, GRPO, LOOP, and GiGPO on AppWorld and AutomationBench while also improving inference efficiency.
Figure 1 : Overview of MoE routing control for agentic trajectories. Top: In each turn of interaction with the environment, agents generate thinking and tool-use fields, then receive environment feedback . Our routing control framework aims to align expert selection with agentic operations (e.g., READ , CREATE , and UPDATE ) in each turn, assigning distinct expert groups to different operations and encouraging similar expert selections within each turn. Bottom: With Qwen3-30B-A3B on AppWorld, our routing control framework (when used with standard RL) unlocks the performance limits and raises the inference throughput, by improving expert specialization (larger with-cross gap of expert usage similarity in Appendix A.1 ) and consistency (higher Jaccard scores).
arXivInterpretation
Off-the-shelf MoE routers already exhibit operation-grouped specialization aligned with agentic trajectories: adjacent tokens within a field tend to reuse experts, and turns with semantically similar operations (e.g., READ, UPDATE) share more similar expert distributions, while other grouping criteria such as tool names do not show this pattern. Prior work largely treated MoE routing as a noise source to be stabilized, or specialized experts by modality or phase; this work is the first to treat the operation semantics of agentic turns as both an observed and an optimized routing structure, with token-level and turn-level statistical evidence. Based on analysis of default routing statistics in off-the-shelf MoE models (Fig. 2, Fig. 8), measured by the gap between within-group and cross-group expert-usage similarity and by within-field Jaccard similarity, with similar patterns observed on Qwen3-30B-A3B and Qwen3.5-35B-A3B.
A hierarchical routing control framework is proposed: turn-level control maximizes mutual information between expert distributions and operation labels to strengthen operation specialization, while token-level control reinforces the previous token's expert set only when adjacent routes are already similar (with a gap threshold) to maintain local consistency; both inject auxiliary gradients at router logits without changing the forward computation or the top-k routing rule. Compared with variants that directly penalize expert switching (switch_reg) or enforce field consensus (routing_reg, outlier_reg), this selective control preserves necessary routing changes at semantic and field boundaries; compared with load-balancing losses alone, it additionally rewards operation-dependent expert usage using operation labels. Ablations on AppWorld with GRPO show token-level control raises within-field Jaccard similarity by 7.7% and turn-level control raises late-stage mutual information by 23.5%, with operation labels outperforming random and application-name labels (Table 6); the outlier_reg variant crashed and switch_reg and routing_reg both underperformed the final framework (Table 2).
An entropy-gated mechanism applies routing-control gradients only when the rollout batch's average policy entropy stays within a target range, detaching them otherwise while standard RL updates continue, thereby avoiding training collapse caused by routing control. Prior MoE RL stability work targeted single-turn tasks or aligned training and inference routing via rollout routing replay (R3); this work uses policy entropy as a switch for auxiliary routing gradients in multi-turn agentic settings and reports curves for no gate, a looser gate, and the target gate. Training curves show that without entropy gating policy entropy spikes and both evaluation splits collapse to zero around step 100; a looser threshold delays collapse to around step 140; with the target threshold training remains stable and retains task performance on both splits (Fig. 4).
End-to-end experiments show the framework consistently improves task performance across multiple RL algorithms and two benchmarks while improving inference efficiency: up to 12.9-point gains on AppWorld and an 11.34-point gain for LOOP on AutomationBench; in steps 100-200 throughput rises from 224 to 316 tokens/GPU/s (+41.1%) and wall time falls from 224 to 174 seconds (-22.3%). Compared with standard agentic RL that only optimizes the policy gradient, this work treats routing structure as an optimizable training signal, improving both success rate and inference efficiency under fixed model, rewards, and RL hyperparameters, and remaining compatible with PPO, GRPO, LOOP, and GiGPO. Trained on 90 AppWorld tasks and 480 AutomationBench public training tasks with Qwen3-30B-A3B-2507 and Qwen3.5-35B-A3B, reporting TGC/SGC along with throughput and wall time; PPO's improvement on AppWorld is inconsistent, which the authors attribute to the small number of trajectories per prompt making mutual information estimation unstable.
Perspective
The framework targets multi-turn agentic RL post-training of sparse MoE models with standard top-k routing during decoding, in environments where tool calls are parseable and operation labels can be defined from them (e.g., AppWorld's Python API calls and AutomationBench's structured tool calls). It does not modify MoE architecture and can be combined with policy-gradient algorithms such as PPO, GRPO, LOOP, and GiGPO, so it can be plugged into existing MoE agent training pipelines as an auxiliary loss. The efficiency mechanism analysis points to batched decoding with batch size greater than 1, where more concentrated expert usage can improve expert-weight reuse and grouped expert computation.
Operation labels are produced by task-specific hard-coded rules, so how to design label taxonomies for new tool ecosystems and how rule errors affect routing control remain open questions. The entropy-gate threshold range and the token-level gap threshold are chosen by search in this work, and their robust ranges across model scales, RL algorithms, and task distributions need further characterization. The efficiency mechanism analysis uses a synthetic diagnostic that reduces the rank of router weight matrices rather than a quality-preserving inference configuration, so the benefit boundary under batched decoding still needs validation in real deployments. In addition, PPO's inconsistent improvement on AppWorld is attributed by the authors to the small number of trajectories per prompt making mutual information estimation unstable, suggesting the mutual information signal's dependence on the number of turns per batch deserves further study.
