Skip to main content
Back to timeline
arXivSource publication:

First joint scaling law for looped mixture of experts: sparsity raises and delays the saturation of recurrence's effective-parameter gain

Synopsis

The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.

AI-generated editorial illustration: Scaling Laws for Looped Mixture of Experts

Interpretation

The paper proposes Loop Scaling Laws, which jointly model recurrence and MoE sparsity (varied through expert count) together with model size and training data in a single functional form, with a sparsity-conditional recurrence mapping: each recurrent pass adds diminishing effective parameters, the total gain approaches a finite asymptote, and sparsity both raises that asymptote and stretches it over more passes. Prior scaling laws treated recurrence or sparsity in isolation: Parcae used a linear, unbounded unrolled parameter count, Iso-Depth used a power-law, unbounded mapping, and MoE laws omitted recurrence, while concurrent looped MoE studies fixed recurrence at two passes in their primary scaling analyses. This work puts both axes into one form and shows it recovers the dense looped law at one pass, the standard MoE law at one pass, and the standard dense law when both are one. The authors sweep jointly over model size, training tokens, expert count, and recurrence, fit linear, power-law, bounded, and sparsity-conditional mappings with the same procedure, and report held-out RMSE: for the dense case the bounded mapping reaches 0.0092 versus 0.0313 for power law and 0.2566 for linear; for MoE the sparsity-conditional bounded mapping attains the lowest RMSE across held-out slices (e.g., 0.0100, 0.0047, 0.0050, 0.0043). Fitting uses more than 1,000 loss observations, and 100 bootstrap refits give 90% confidence intervals for coefficients and for the joint optimum.

The fitted law is used as a design tool: it selects the compute-optimal recurrence at fixed sparsity, the memory-optimal expert count at fixed recurrence, and the joint optimum under given training-compute and weight-memory budgets. This moves the scaling law from loss prediction to architecture selection under resource constraints, and makes explicit two complementary trade-offs: at fixed weight memory, recurrence trades extra compute for quality; at fixed compute, sparsity trades extra weight memory for quality. The paper gives three selection criteria (Eqs. 11, 12, 13), uses the fitted-law RMSE as the minimum distinguishable loss improvement, and searches 300 candidate configurations spanning five model scales, several expert counts, and several recurrences. Reported trends include compute-optimal recurrence rising with the training-compute budget, smaller active models favoring higher recurrence, greater sparsity shifting the compute-optimal regime toward higher recurrence, and tighter memory favoring higher recurrence with smaller expert counts (e.g., a 1.0B active model with recurrence 3 and 8 experts at 3 GB; a 1.6B active model with recurrence 2 and 16 experts at 10 GB).

Across 14 downstream benchmarks, sparsity and recurrence deliver complementary gains: sparsity yields active-parameter efficiency, recurrence yields total-parameter efficiency with a reasoning tilt, and scaling both together pushes the performance frontier further. This supplies downstream-task evidence for the complementarity of the two axes, beyond loss-prediction fits. Evaluation covers 14 benchmarks in five categories: reasoning (BBH, GSM8K), science (ARC-C/E, OpenBookQA), commonsense (HellaSwag, PIQA, SIQA, WinoGrande), reading (BoolQ, DROP), and knowledge (MMLU, Natural Questions, TriviaQA), comparing 0.3B, 0.6B, and 1.0B active sizes under varying expert count and recurrence. The paper reports that at matched training FLOPs a 0.3B MoE can surpass a larger 1.0B dense model, that recurrence gains can offset a doubling of total parameters, and that models increasing both sparsity and recurrence consistently outperform the dense baseline.

In a trillion-token case study, a 0.3B-active/1.3B-total looped MoE trained with law-derived recurrence matches a 0.6B-active/2.9B-total non-looped MoE on reasoning benchmarks at matched training compute, and can scale test-time compute on demand by varying recurrence at inference. This carries the law's predictive power to a realistic training scale and shows that extra inference compute from recurrence can be traded for parameter efficiency, while providing an inference-time compute knob. The two models are trained at trillion-token scale (the looped MoE on fewer tokens than the non-looped MoE) at approximately matched training FLOPs; Table 2 shows the looped MoE matching the larger non-looped MoE on BBH and GSM8K while trailing on Overall, and raising inference recurrence from 1 to 5 substantially improves Overall.

Perspective

The law targets designing looped MoE models under given training-compute and weight-memory budgets, and applies to the authors' model ladder, data recipe, and middle-block looping strategy; it states asymptotic properties within the observed recurrence range rather than guarantees beyond it. For a reader, it offers an actionable recipe: fit the law, then pick compute-optimal recurrence via Eq. 11, memory-optimal expert count via Eq. 12, and the joint optimum via Eq. 13, trading parameters, compute, and memory in resource-constrained deployment. The paper also points to future directions, including richer looping strategies, broader scales from edge models to cloud models, and reducing looped-model memory and latency by optimizing recurrent states or adding early-exit gating.

The asymptotic bound is a property of the mapping rather than a guarantee beyond the observed recurrence range, so recurrence and expert counts outside the fitted region warrant caution; memory optimization covers weight memory only, leaving KV cache and other runtime state out of scope; in downstream evaluation the looped MoE still trails the larger non-looped MoE on Overall and only matches it on reasoning tasks; and because the law is fitted on the authors' own model ladder and data recipe, whether its coefficients remain stable under other looping strategies or data distributions is an open question.

Sources