DEP prunes experts per request from prompts in multi-agent systems, beating static pruning and merging baselines most clearly when few experts are kept
Related research and updatesSynopsis
The work proposes Dynamic Expert Pruning (DEP): its analysis shows that in multi-agent systems different tasks and roles recruit different experts while static pruning applies one fixed mask to every request; DEP uses an agent's system and task prompts to produce a per-request expert mask in a single forward pass with no per-configuration calibration, achieving better overall accuracy than static pruning and merging baselines across diverse tasks, roles, model scales, and MoE architectures, generalizing to workflows unseen in training, with the largest margin when few experts are retained.
Figure 1: Static masks that track the dense model for a single agent lose much of that accuracy on a multi-agent workload. Left: identical pruning methods, backbone, retention ( ρ = 50 % \rho{=}50\% ), and calibration data; only the serving setting changes. Right: the resulting accuracy–memory frontier.
arXivInterpretation
The analysis shows that in multi-agent systems different tasks and roles recruit different experts, whereas static pruning applies a single offline-calibrated expert subset to every request. Prior expert pruning was static, with one mask calibrated offline and reused for every subsequent request; this work identifies multi-agent settings as where that assumption fails under heterogeneous workloads. Based on the authors' analysis of differing expert recruitment across tasks and roles, stated in the abstract as "our analysis shows," without reported statistics.
An agent's system and task prompts are by themselves sufficient to identify the experts that agent and its task require, because that text already describes what the agent will do. It shifts the basis for expert selection from offline calibration or runtime observation to the request's own prompt text, allowing the mask to be determined at request granularity. The abstract presents this as a finding established in the paper ("a finding we establish here"); validation details are not expanded in the abstract.
DEP trains a lightweight predictor once on workflow transcripts so that prompts become a specialized per-request mask in a single forward pass, with no per-configuration calibration. Relative to the offline single-mask pipeline of static pruning, DEP turns mask generation into a one-pass online prediction and avoids recalibration for each configuration. The method description comes from the abstract; training data are workflow transcripts and the predictor is described as lightweight, with no parameter counts or training-cost figures given.
Across diverse tasks and roles, model scales, and MoE architectures, DEP achieves better overall accuracy than static pruning and merging baselines, generalizes to workflows unseen in training without retraining, and has its largest margin over those baselines when few experts are retained. Compared with static pruning and merging baselines, DEP attains higher overall accuracy under heterogeneous multi-agent workloads and indicates that role specialization permits sparser serving than static pruning allows. The abstract reports comparisons across tasks, roles, model scales, and architectures and notes the margin grows as fewer experts are retained; no specific accuracy values or retention ratios are given.
Perspective
The result targets multi-agent workflows whose behavior is described by prompts: one MoE backbone serving many tasks and roles at once, with system and task prompts available. In that setting DEP lets the serving side choose an expert subset per request, sustaining higher overall accuracy when few experts are retained and applying to workflows unseen in training without retraining; for engineering teams deploying MoE multi-agent systems under limited memory, this offers a finer-grained memory-accuracy control than static pruning.
The abstract gives no specific accuracy values, expert retention ratios, predictor size, or training cost, and does not say whether masks remain reliable when prompts are missing, low quality, or vague about the task. Whether the margin is consistent across MoE architectures and model scales, and over what range the trend of a larger margin at higher sparsity holds, still requires the paper's tables and experimental setup. Because this assessment rests on the abstract alone, these details are open questions to verify.
