Public articles linked to the same research event.
arXiv The work proposes Dynamic Expert Pruning (DEP): its analysis shows that in multi-agent systems different tasks and roles recruit different experts while static pruning applies one fixed mask to every request; DEP uses an agent's system and task prompts to produce a per-request expert mask in a single forward pass with no per-configuration calibration, achieving better overall accuracy than static pruning and merging baselines across diverse tasks, roles, model scales, and MoE architectures, generalizing to workflows unseen in training, with the largest margin when few experts are retained.
The work proposes Dynamic Expert Pruning (DEP): its analysis shows that in multi-agent systems different tasks and roles recruit different experts while static pruning applies one fixed mask to every request; DEP uses an agent's system and task prompts to produce a per-request expert mask in a single forward pass with no per-configuration calibration, achieving better overall accuracy than static pruning and merging baselines across diverse tasks, roles, model scales, and MoE architectures, generalizing to workflows unseen in training, with the largest margin when few experts are retained.
The work proposes Dynamic Expert Pruning (DEP): its analysis shows that in multi-agent systems different tasks and roles recruit different experts while static pruning applies one fixed mask to every request; DEP uses an agent's system and task prompts to produce a per-request expert mask in a single forward pass with no per-configuration calibration, achieving better overall accuracy than static pruning and merging baselines across diverse tasks, roles, model scales, and MoE architectures, generalizing to workflows unseen in training, with the largest margin when few experts are retained.
The work proposes Dynamic Expert Pruning (DEP): its analysis shows that in multi-agent systems different tasks and roles recruit different experts while static pruning applies one fixed mask to every request; DEP uses an agent's system and task prompts to produce a per-request expert mask in a single forward pass with no per-configuration calibration, achieving better overall accuracy than static pruning and merging baselines across diverse tasks, roles, model scales, and MoE architectures, generalizing to workflows unseen in training, with the largest margin when few experts are retained.