Skip to main content
Back to timeline
arXivSource publication:

IntBMoE: Decoupling Expert Participation, Execution, and Materialization via Block-Level Conditioning

Synopsis

The work proposes IntBMoE, a block-conditioned Mixture-of-Experts architecture in which a small learned codebook defines reusable multi-layer blocks, a shared hypernetwork merges all expert bases of each layer into one composed expert, and a router executes only a few blocks per token, thereby keeping pool-wide participation while retaining sparse execution and bounded parameter materialization; it reaches 73.76% Top-1 and 91.48% Top-5 on ImageNet-1K, outperforms the compared sparse and dense MoE baselines on MiniPile language modeling and IntTravel sequential recommendation, and is fully deployed in AMap's generative recommendation system with a 2.4% relative UVCTR gain in online A/B testing.

AI-generated editorial illustration: IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

Interpretation

It decouples the three quantities of a Mixture-of-Experts layer — participation, execution, and materialization — so they can be set independently. Existing designs couple them: sparse routing keeps execution and materialization low but shrinks participation, since "non-selected experts neither affect the token's output nor receive a learning signal from it"; dense output mixing restores full participation but its execution grows with the number of experts; parameter merging keeps execution at one expert but its materialization grows with the number of routing decisions. IntBMoE lets the full pool first construct a bounded set of reusable transformations, after which each token independently selects and executes only a few of them. The paper provides a complexity analysis: composing all blocks costs on the order of O(N·L·d²) and executing the selected blocks O(N·k·L·d²), with the composition term paid once for reusable blocks; on ImageNet-1K it reaches 73.76% Top-1 with about 24M total parameters and 23.111M activated parameters.

A codebook plus a shared hypernetwork generates "composition recipes", making block parameters input-independent and therefore precomputable and cacheable. Unlike SMEAR and Lory, which merge parameters per token or per segment, IntBMoE conditions on a finite, input-independent codebook embedding, and the hypernetwork produces compact composition coefficients rather than directly generating full weight matrices, so composed blocks can be reused across routing decisions; the authors describe it as "a basis-constrained hypernetwork with bounded parameter materialization". With caching, inference cost drops from 4.063 GFLOPs to 3.457 GFLOPs; as the expert pool grows from 1 to 128, uncached peak memory and compute rise, whereas cached versions stay at a fixed value and fixed GFLOPs, and caching becomes more memory-efficient from 16 experts onward.

Dual-Path Residual Gating (DPRG) couples two independently composed paths through multiplicative gating, adding expressiveness without enlarging the expert pool. Each path is still composed linearly from expert bases, but DPRG introduces residual multiplicative modulation between the value and gate paths, making the overall transformation nonlinear in those bases; ablations show removing the gate path hurts most (ImageNet-1K Top-1 falls from 0.7376 to 0.6804, MiniPile PPL rises from 14.5878 to 15.8947). Component ablations on ImageNet-1K, MiniPile, and IntTravel consistently show the gate path has the largest impact, and the parameter-matched single-layer block variant also underperforms the full two-layer block.

Full-pool participation is shown to be effective, and blocks do learn distinct composition recipes. Rather than only asserting pool-wide participation, the authors remove expert bases layer by layer: across 64 removal settings in Layers 0, 2, 4, and 6, removing any examined expert basis lowers Top-1, with the smallest drop at 0.26 percentage points; meanwhile different blocks in the same layer assign different positive and negative weights to the same expert bases, and recipe differentiation grows with depth. The removal study evaluates the original checkpoint without retraining and compensates for the scale change in composed weights; mean pairwise cosine similarity between value-path block recipes falls from 0.796 at Layer 0 to 0.010 at Layer 6, and gate-path similarity from 0.726 to 0.043.

Perspective

The result targets model designers who need to expand expert capacity under limited compute and memory, and applies to settings such as visual classification, autoregressive language modeling, and sequential recommendation; the caching mechanism assumes model parameters are fixed at inference time so that block synthesis leaves the request path, which means its benefits differ between training and inference. For latency-sensitive production systems, the stated conditions are a fixed block configuration and precomputed caching, under which the AMap deployment meets the 60ms budget.

A careful reader would still watch how the growing recipe differentiation with depth behaves under larger backbones and more MoE layers; the expert-basis removal study covers only the 16 bases in Layers 0, 2, 4, and 6, so contribution distributions in other layers and larger pools remain open; the online A/B test spans one week, leaving long-term effects and transferability to other business settings as open questions; and since this evidence bundle is a full-text parse, verifying specific figure values and appendix details still calls for the original paper.

Sources