Shared Low-rank Basis Factorization (SLBF) compresses MoE expert weights data-free and beats pruning, merging, and reconstruction baselines across five MoE architectures from 16B to 122B
Synopsis
The authors analyze three MoE compression families (expert pruning, expert merging, weight reconstruction) and derive structural error bounds showing that pruning and merging incur non-vanishing errors tied to routing and expert heterogeneity, whereas weight reconstruction avoids these structural costs by preserving expert structure and routing; motivated by this, they propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that replaces full-rank shared bases with rank-r bases shared among experts and applies a bilinear gauge fixing to remove redundant parameters, consistently outperforming methods from all three families across five MoE architectures spanning 16B to 122B parameters.
(a) Pair-level structural costs.
arXivInterpretation
The paper provides a unified structural cost analysis of the three MoE compression interventions and derives new error bounds for expert pruning and weight reconstruction, showing that pruning and merging carry a structural term induced by expert removal or tied routing, while weight reconstruction error is controlled solely by the parameter reconstruction residual. Prior work largely compared compression methods empirically; this work places pruning, merging, and reconstruction in one pairwise analysis framework with comparable error expressions and states when the structural term vanishes (static routing on the pair or functionally identical experts). Derivations are carried out for a pair of experts under a weak-correlation approximation and extended to SwiGLU-style FFNs, with a prefactor set by original weight norms and the activation's Lipschitz constant; on Qwen3-30B-A3B and Mixtral-8x7B, C4 activation statistics show both structural quantities strictly above zero in every MoE layer and growing 3-4 orders of magnitude with depth.
The paper proposes Shared Low-rank Basis Factorization (SLBF), replacing each full-rank shared basis in MoBE with a rank-r factorization so that more bases fit under a fixed compression budget, letting experts compose across multiple bases rather than concentrating on a single dominant one. Relative to MoBE's few full-rank bases, SLBF trades few-rich bases for many-simple bases with richer cross-expert sharing; the authors note the pre-activation mixture of low-rank bases generically attains the same per-basis rank as MoBE when r is full, so the substitution does not lower the per-basis expressiveness ceiling. Under matched budget SLBF converges faster and reaches lower reconstruction MSE, uniformly across all 32 MoE layers of Mixtral-8x7B; a sweep on Qwen3-30B layer 4 gate_proj places the optimum near a basis count matching the number of experts, which the authors adopt.
The paper proposes a bilinear symmetry gauge fixing that removes redundant parameters per basis at no representational cost. The step reinvests freed parameters as higher per-basis rank, an engineering use of the known bilinear factorization symmetry rather than a change of model form. Ablation shows SLBF without gauge fixing already outperforms MoBE at every compression ratio, with gauge fixing adding a consistent secondary gain that widens at more aggressive compression; the authors report the full-column-rank condition held for all trained factors in their experiments.
Evaluated on five MoE architectures (16B to 122B) in a one-shot, data-free setting with no post-compression fine-tuning, SLBF consistently outperforms representative methods from all three compression families on reasoning, knowledge, math, and coding benchmarks. Against same-family MoLAE and MoBE, SLBF leads even at matched projection scope and gains further by distributing the budget across all three FFN projections; against pruning and merging methods its advantage widens at higher compression and on more architectures. On Qwen3-30B at 24% compression SLBF averages 84.3 versus 82.8 for MoBE while pruning/merging baselines drop over 8 points; the gap over MoBE reaches 14.3 points on Gemma4-26B and 5.2 points on Qwen3.5-122B; Moonlight-16B at 15% compression drops only 1.8 points versus 9.9 for MoBE and over 30 for MoLAE; Mixtral-8x7B at 24% compression reaches 70.7 average and 6.92 WikiText-2 perplexity, beating the strongest pruning baseline EAN (70.2/8.28) and the strongest merging baseline HC-SMoE (69.1/8.54).
Perspective
The result targets engineering and research settings that must deploy MoE large language models under limited accelerator memory or distribution bandwidth, especially one-shot compression pipelines that want to preserve the sparse routing structure without calibration data or post-compression fine-tuning. Because the method fits each layer and projection independently in weight space, it is naturally parallel and suits shard-wise processing of large checkpoints. The authors report that in runtime-decompressed serving on Mixtral-8x7B, on-device model weights drop by 30%, the freed memory becomes KV-cache budget, and maximum concurrent requests rise by 53% while throughput stays close to the uncompressed baseline, indicating a workable path on the inference-serving side. Applicability is bounded by the architectures and compression ratios evaluated: five MoE architectures from 16B to 122B parameters at roughly 15% to 37% total-model compression, with Mixtral compressing only gate and up projections while down_proj stays full rank.
The structural error bounds rest on a pairwise-expert, weak-correlation approximation with linear or SwiGLU experts, so how tight they remain under more complex routing and larger expert counts is an open question. Evaluation is a single run; the authors state the initialization is deterministic so results are reproducible under fixed hyperparameters, but run-to-run variance is not reported. Because compression is data-free with no post-compression fine-tuning, calibration-guided budget allocation and fine-tuning to recover residual accuracy loss are not explored, nor are interactions with other compression axes such as quantization or attention compression. The decomposition structure is fixed across models and layers, and adaptive per-layer or per-projection choices of the basis count may yield further improvements. The deployment evaluation covers limited hardware and serving configurations, runtime-decompressed inference shows a memory-throughput trade-off, and directly loading the factorized representation can incur substantial reconstruction overhead for large MoEs in the current implementation. In addition, the original text carries key derivations and curves in equations and figures; in this parse those equations and figure captions are not visible, so the exact form of the error bounds and the numerical details of Figures 1-5 are summarized from the surrounding prose, and readers needing precise expressions should consult the original appendix.
