Skip to main content
Back to timeline
arXivSource publication:

SplitMoE splits the video-diffusion expert pool into semantic and generic branches, beating a same-source MoE at 14B activated parameters and showing coarse-to-fine denoising routing

Synopsis

The work identifies a "uniformity trap" in which language-style token-wise MoE routing, combined with uniform expert-usage regularization, scatters neighboring patches of the same object across different experts in video diffusion, causing routing fragmentation and structural distortion; it proposes SplitMoE, which bifurcates the expert pool into semantic and generic experts and uses prototype-guided routing with pull-push regularization so tokens cluster by semantic attributes, outperforming traditional load-balanced MoE on VBench-2.0 and T2V-CompBench under an equivalent activated-parameter budget, while revealing a coarse-to-fine denoising pattern in which semantic-expert weight peaks in the early high-noise stage around steps 10-15 and then shifts toward generic experts.

AI-generated editorial illustration: Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Interpretation

The paper frames the "uniformity trap" as a core problem for video MoE: standard MoE inherits load balancing from language modeling and spreads tokens uniformly across the expert pool, whereas video tokens are spatiotemporally redundant and semantically long-tailed, so tokens from the same semantic region get dispersed. Prior visual-generation MoE work largely adopted language-centric designs or decoupled experts by conditioning type or task modality (e.g., ProMoE, MammothModa2); this work locates the problem in the spatiotemporal redundancy and semantic imbalance of video structure itself. Video tokens are grouped with off-the-shelf semantic segmentation masks and two metrics are measured, semantic distinctiveness and intra-sample routing dominance; standard MoE shows consistently low values across most semantic classes, and Top-1 expert maps illustrate temporal jitter, spatial striping, and semantic fragmentation.

SplitMoE explicitly partitions the expert pool into two disjoint subsets, semantic experts for high-level semantic abstraction and generic experts for residual visual information and flexible generative capacity, using independent sigmoid scores so one token can be highly compatible with both a semantic and a generic expert. Unlike a single homogeneous pool with global softmax competition, this role-aware split lets the semantic and generic branches serve different functions instead of jointly optimizing low-frequency semantics and high-frequency detail in one pool. Built by upcycling the low-noise Wan 2.2 checkpoint, replacing dense FFNs at even-numbered layers from 15 to 35 with sparse MoE layers; each MoE layer has one continuously active shared expert and 100 routed experts (20 semantic, 80 generic), and Top-k routing activates 14B of 27B total parameters per forward pass, matching the dense baseline's compute.

Prototype-guided semantic routing uses learnable prototypes as a bridge between the VAE feature space and DiT tokens: the router alignment target comes from temperature-scaled cosine similarity between clean VAE features and prototypes, while the prototypes themselves are optimized by a pull term toward current-batch and historical features and a push term with an inter-prototype margin. Unlike ProMoE, which performs prototypical matching directly within DiT layers, this work uses prototypes as a bridge between DiT router logits and reconstructive VAE features, and shows pull and push form a coupled mechanism. Ablations show the variant keeping the 20/80 partition but removing prototype guidance performs similarly to standard MoE, indicating partitioning alone brings limited gains; variants keeping only pull or only push are unstable on quantitative metrics and can even underperform the no-guidance variant, with worse validation diffusion loss; VAE-space visualization shows that without push, prototypes are attracted to dominant visual regions and cover the manifold poorly, while without pull, prototypes drift away from the data manifold and produce dead semantic experts.

Under matched activated-parameter budgets, full SplitMoE outperforms the same-source dense fine-tuning and standard MoE on VBench-2.0 and T2V-CompBench, and exhibits a coarse-to-fine routing pattern over denoising. The paper emphasizes that gains come from role-aware capacity allocation rather than a larger compute budget, and reports stage-dependent routing that emerges without any explicit timestep-conditioned routing constraint. In same-source comparisons, Dense Wan2.2-FT is 14B while Wan2.2-MoE and SplitMoE variants are A14B; full SplitMoE reaches VBench-2 Creativity 58.46%, Common Sense 64.89%, Human Fidelity 84.47%, Physics 69.41%, and T2V-CompBench Consist attr. 84.63%, Numeracy 44.01%; validation diffusion loss shows comparable loss at roughly 70% of training steps; over a 50-step denoising trajectory, semantic-expert weight peaks around steps 10-15 and then declines.

Perspective

The result targets text-to-video generation with diffusion-transformer backbones, realized by MoE upcycling on the Wan 2.2 low-noise checkpoint, and is meant for research and engineering teams that want to improve generation quality and convergence speed without increasing the activated-parameter budget; the routing metrics, Top-1 expert maps, and denoising-trajectory analysis can serve as reference tools for diagnosing video MoE routing, and the coarse-to-fine temporal pattern offers a design basis for allocating expert capacity by stage.

The limitations stated in the paper include reliance on frozen visual features and the overhead of sparse routing, pointing toward adaptive partitioning, efficient distributed routing, and broader video generation tasks; in addition, cross-model comparisons involve differences in training data, scale, compute budget, and post-training recipes, which is why the paper centers its conclusions on same-source ablations, and readers should keep that scope in mind when interpreting comparisons with independently developed models. Pull and push must be used together in prototype guidance, since either term alone may introduce a biased routing prior, and whether this coupling remains stable when transferred to new data distributions is an open question worth watching.

Sources