Public articles linked to the same research event.
arXiv This work systematically studies how agentic post-training interacts with MoE expert selection, observing that off-the-shelf MoE routers already show operation-grouped specialization (e.g., READ, UPDATE) that standard RL perturbs; the authors propose a hierarchical routing control framework that aligns turn-level expert selection with operation labels via mutual information, regularizes token-level adjacent consistency with a gap threshold, and adds entropy gating for stability, improving success rates by over 10 points across PPO, GRPO, LOOP, and GiGPO on AppWorld and AutomationBench while also improving inference efficiency.
This work systematically studies how agentic post-training interacts with MoE expert selection, observing that off-the-shelf MoE routers already show operation-grouped specialization (e.g., READ, UPDATE) that standard RL perturbs; the authors propose a hierarchical routing control framework that aligns turn-level expert selection with operation labels via mutual information, regularizes token-level adjacent consistency with a gap threshold, and adds entropy gating for stability, improving success rates by over 10 points across PPO, GRPO, LOOP, and GiGPO on AppWorld and AutomationBench while also improving inference efficiency.
This work systematically studies how agentic post-training interacts with MoE expert selection, observing that off-the-shelf MoE routers already show operation-grouped specialization (e.g., READ, UPDATE) that standard RL perturbs; the authors propose a hierarchical routing control framework that aligns turn-level expert selection with operation labels via mutual information, regularizes token-level adjacent consistency with a gap threshold, and adds entropy gating for stability, improving success rates by over 10 points across PPO, GRPO, LOOP, and GiGPO on AppWorld and AutomationBench while also improving inference efficiency.
This work systematically studies how agentic post-training interacts with MoE expert selection, observing that off-the-shelf MoE routers already show operation-grouped specialization (e.g., READ, UPDATE) that standard RL perturbs; the authors propose a hierarchical routing control framework that aligns turn-level expert selection with operation labels via mutual information, regularizes token-level adjacent consistency with a gap threshold, and adds entropy gating for stability, improving success rates by over 10 points across PPO, GRPO, LOOP, and GiGPO on AppWorld and AutomationBench while also improving inference efficiency.