Skip to main content
Back to timeline
arXivSource publication:

Higher-Order Pruning of Experts in Mixture-of-Experts Language Models

Synopsis

The work introduces HOPE, a second-order expert-pruning objective that records a per-layer pairwise expert interaction matrix F during calibration and solves a quadratic program to select the prune-set, provably showing that first-order methods such as REAP are special cases that ignore interaction terms; across three frontier MoE models (up to 122B parameters), two calibration sets, and multiple benchmarks, HOPE achieves the best average rank in most conditions, with an average rank of 1.58 at 50% pruning (versus 2.42 for the next-best method, REAP) and gains of up to +6.1% on agentic coding.

Source-provided article image: Higher-order pruning of experts in mixture-of-experts language models
Figure 2 ·

Figure 2: The interaction matrix F F recorded by HOPE exhibits expert cooperation. a) F F -matrix for GLM-4.5-Air at three representative layers (early, middle, late). Experts have been hierarchically clustered per layer for visualization. Visible block structure indicates groups of experts with high pairwise interaction scores. b) Distribution of diagonal vs off-diagonal terms in F F across all layers (GLM-4.5-Air). Off-diagonal (interactive) terms are substantial in magnitude (compared to the diagonal), showing that pairwise interactions are a non-negligible component of the pruning objective.

arXiv

Interpretation

It proposes HOPE, a second-order pruning objective that formulates expert pruning as a quadratic program p⊤Fp over an interaction matrix F, where diagonal entries capture individual expert contributions and off-diagonal entries capture pairwise co-contributions. Prior first-order methods (Frequency, EAN, REAP, MAN) assign each expert an independent scalar importance score and ignore interactions between experts; HOPE is the first to incorporate pairwise interactions into pruning decisions. The paper provides a full derivation (Appendix A): it decomposes pruning error into substitution and renormalization error, proves the squared substitution error is bounded by (1+ρ)²·Z, and proves that minimizing E[Z] is equivalent to solving the quadratic program (Theorems 1 and 2).

It shows REAP is a special case of HOPE when interaction terms are ignored: the diagonal F_k,k=E[(g_k‖f_k‖)²] differs from the squared REAP score only by the variance of expert contributions, and zeroing off-diagonals empirically recovers the same prune-set as REAP. It places the strongest existing first-order method inside the new framework, so any improvement of HOPE over REAP is directly attributable to the off-diagonal interaction terms. Beyond the derivation, the paper reports a near-perfect Spearman correlation (ρ=0.988, Figure S4) between the square root of F_k,k and the REAP score.

Across 3 MoE models (Qwen3.5-122B-A10B, Qwen3.5-35B-A3B, GLM-4.5-Air), 2 calibration sets, 6 pruning rates, and the Tulu3 and SWE-Bench Pro benchmarks (54 configurations), HOPE has the best overall average rank (2.07) and an average head-to-head win rate of 73%; its advantage grows with pruning rate, reaching an average rank of 1.58 at 40–50%. Prior methods had not been systematically compared against a second-order approach across this many models, calibration sets, and pruning rates; HOPE's largest advantage lies in the high-pruning-rate regime most relevant to deployment. Each condition reports mean and standard deviation over 3 independent calibration trials (Tables S1–S6); at 50% pruning HOPE leads REAP on agentic coding by +2.8% on average and up to +6.1%; HOPE finishes last in only 1/54 conditions (gap -1.7%), versus 5 for REAP and 43 for Frequency.

Analysis shows the off-diagonal entries of F are not noise: their average magnitude is about 33% of the diagonal and they exhibit block structure; when HOPE disagrees with a first-order method, HOPE preferentially retains experts with stronger interactions with other surviving experts. It provides mechanistic evidence for why second-order pruning works, indicating first-order methods are misled by experts' self-importance and ignore cooperative structure. Based on hierarchical clustering visualizations of F across layers of GLM-4.5-Air and diagonal/off-diagonal distribution statistics (Figure 2), plus cooperative-score comparisons of experts classified by disagreement (Figure 3); prune-set disagreement is distributed across all layers (Figure 4b).

Perspective

The results target top-K routed MoE language models (128–256 experts with top-8 routing in the experiments) and are validated on coding, math, instruction-following, and agentic coding benchmarks; the method prunes each layer independently, shared experts are never pruned, and it composes with quantization. For deployers seeking longer context, more concurrent requests, or longer reasoning traces on the same hardware, HOPE offers a pruning scheme that holds up at high pruning rates; the paper also presents a theoretical framework and preliminary results for a cross-layer extension, CHOPE, to explore non-uniform per-layer pruning budgets.

HOPE currently models only up to second-order (pairwise) interactions, and higher-order interactions are combinatorially difficult to quantify; preliminary results for the cross-layer extension CHOPE show that its non-uniform budget allocation can introduce instability, and the paper suggests adding per-layer minimum-expert constraints; moreover, prune-set agreement across different calibration domains is lower than across repeated trials within one domain (HOPE's average calibration Jaccard is 0.60 versus REAP's 0.61), so the choice of calibration data remains a variable for readers to watch.

Sources