Skip to main content
Back to timeline
arXivSource publication:

E-MoE uses expert routing as a discrete latent, cutting LM1B one-step generative perplexity from 1433.8 to 643.8

Synopsis

The work proposes E-MoE, which treats the expert-routing decisions of a Mixture-of-Experts backbone as a discrete shared latent so that the masked diffusion reverse process becomes a mixture of factorized distributions, improving few-step generation without increasing active parameters per token: it beats factorized baselines on synthetic 2-D densities, binarized MNIST and LM1B, lowering LM1B generative perplexity at NFE=1 from MDLM's 1433.8 to 643.8.

AI-generated editorial illustration: E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

Interpretation

E-MoE writes the reverse process as a mixture of factorized distributions conditioned on a discrete shared latent, which is exactly the per-token, per-layer expert-routing decision the MoE backbone already computes. Earlier continuous-latent routes such as VADD need an extra recognition network and a fixed Gaussian prior; E-MoE reuses the same router, evaluated on the noisy and the clean sequence, to supply prior and posterior, so no recognition model is added and no prior has to be designed. The paper derives an upper bound on the negative ELBO for this mixture reverse process and reduces the latent KL to a per-layer, per-token categorical quantity computed from the router itself; training handles the discrete route with a Gumbel-Softmax straight-through estimator.

In the few-step regime where the factorization error binds, E-MoE generates markedly better than both the factorized baseline and the continuous-latent one. The paper reports that at NFE=1 and NFE=2 E-MoE lowers generative perplexity by roughly 50% relative to every baseline at comparable sample entropy (about 4.34-4.35), and that MAUVE separates the models even more at NFE=1. On LM1B all three models share the same small DiT backbone, optimizer and training budget, and E-MoE uses as many active parameters per token as MDLM; results are reported with generative perplexity, sample entropy and MAUVE.

On two-dimensional synthetic densities and binarized MNIST, E-MoE recovers cross-position structure that factorized sampling loses. At low NFE the factorized MDLM combines coordinates from different clusters into spurious modes or spreads mass around the manifold; E-MoE keeps samples on the eight true clusters and the spiral, and reaches the best test BPD among the three on binarized MNIST (0.062). The toy task measures validity, the share of samples landing on the data support, where E-MoE exceeds VADD on 8-modes from NFE=1; on MNIST E-MoE has 2.49M parameters against VADD's 2.50M.

An ablation shows the few-step gain comes from the mixture itself rather than parameter count: making routing greedy turns the latent into a point mass and performance falls back toward MDLM. On the same checkpoint with the same sampler, switching the mixture off raises generative perplexity from 626.1 to 1079.5 at NFE=1 and from 389.0 to 834.2 at NFE=2, while sample entropy stays essentially unchanged. The ablation changes only how the route is sampled, with weights held fixed; the paper also reports that E-MoE's routing KL settles at about 2 nats per token without warmup, whereas VADD's latent KL falls to about 0.1 nats per token and stays there.

Perspective

The result targets low-NFE few-step generation, the setting where each reverse step reveals many tokens at once; as the number of denoising steps grows, the benefit naturally diminishes. It applies to backbones that already have expert routing, and the paper points specifically to LLaDA-style and LLaDA-MoE architectures, where routing decisions already exist and only need to be reinterpreted as discrete shared latents and trained to coordinate multi-token prediction. Experiments cover two-dimensional synthetic densities, binarized MNIST and unconditional LM1B generation, all compared under matched backbone, optimizer and training budget.

The paper itself notes the method relies on meaningful router alignment: if the router on the noisy sequence cannot recover the clean routing decisions, or if the routing distribution collapses, the mixture degenerates toward a single factorized component. The gain narrows as NFE grows, so its value is concentrated in the few-step regime. In addition, several appendix tables and figures in the loaded text (such as the exact MAUVE value table, generated samples and training curves) appear as placeholders, so their specific numbers and sample details cannot be checked here; these are open questions to confirm against the original.

Sources