Skip to main content
Back to timeline
arXivSource publication:

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

Synopsis

Using a unified prelude–core–coda looped-transformer family, this work fits a separate compute-optimal recipe for each architecture and trains scaling ladders to about 10^20 FLOPs on FineWeb, finding that model growth (raising the number of core passes mid-training) and a boundary operator (normalizing the residual stream and re-injecting the prelude output) can change pre-training scaling exponents so that compute-efficiency gains widen with scale, while untying weights moves only the constant; in data-constrained multi-epoch training the optimal loop count grows with compute and scaling loops is more compute-efficient than scaling model size.

Source-provided article image: How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
Figure 5 ·

Figure 5: In multi-epoch training, the optimal loop count grows with compute, scaling the loop count beats scaling model size even against tuned weight decay, and looping leaves the optimal weight decay nearly unchanged. Models are trained on 100M unique FineWeb tokens for 10 epochs, and all losses shown are validation losses on FineWeb. First: matched-compute cuts. The loss of each loop count is interpolated in compute between neighbouring depths, curves are quadratics in log loss against log loop count, and stars mark their minima. The optimal loop count rises from 1.4 to 6.7 across budgets, the gap to Operator-1 grows from 0.00 to 0.11. Second: marginal gain from looping. For each depth, we plot the loss at loop count K K minus the Operator-1 loss at the same depth. Every curve decreases monotonically, and from d ​ 8 d8 upward the curves lie within 0.006 of one another at every loop count, so the gain from looping appears nearly independent of depth. Third: weight decay when scaling loops versus model size. Scaling model size at fixed weight decay overfits, and retuning weight decay at every model size only reaches 3.40 loss at 7.5 × 10 18 7.5\times 10^{18} FLOPs. Scaling the loop count at fixed model size and fixed weight decay matches that loss with 2.2 × 2.2\times less compute, and tuning weight decay on top gains at most 0.02, so loop-count scaling is more compute-efficient and nearly free of tuning. Fourth: optimal weight decay by depth and loop count. The optimum rises with depth but is nearly flat in loop count, so overfitting tracks stored parameters rather than executed depth.

arXiv

Interpretation

Model growth, doubling core passes from 2 to 4 partway through training, improves the scaling exponent rather than only the constant. Earlier growth and looping studies report gains at one or a few target sizes rather than a change in the compute-optimal scaling law; here Untied-Grow's compute multiplier over Vanilla widens from 1.30× at 10^18 FLOPs to 1.55× at 10^20 FLOPs, and its multiplier over Untied-2 grows from 1.09× to 1.16×. Eight compute-optimal ladders with a separately fitted recipe per architecture, log–log slope regression standard error below 10^-3; growth-timing sweeps show interior minima at fixed size and budget, and grown models prefer a smaller starting model with more tokens (optimal tokens per parameter rising from 6 to 7–8).

Adding only the boundary operator, Norm(h)+αe applied before each core pass and before the coda, improves the scaling exponent. Prior depth-related interventions improve loss at fixed model size but had not been shown to change the scaling exponent; here Operator-1 matches Vanilla in parameters and FLOPs yet its multiplier rises from 1.12× at 10^18 FLOPs to 1.25× at 10^20 FLOPs, and ablations show that removing normalization, injection, or coda injection reduces the gain. Parameter- and FLOPs-matched comparisons plus ablations of the operator's components (Table 4, Figure 15(a)); KL effective depth reaches 24 for Untied-2 versus 20 for Deep Vanilla at 10^20 FLOPs, and width-only scaling turns the exponent advantage into a constant one.

Untying weights yields only a fixed constant-factor improvement and does not change the scaling exponent. This separates the effect of depth from the effect of weight sharing: Untied-2 and Loop-2 share the same computation graph and FLOPs, and Untied-2 sits 1.08× above Loop-2 at 10^18 FLOPs and 1.06× at 10^20 FLOPs, a factor that does not widen with scale; tying weights while growing therefore retains the exponent improvement of growth at Vanilla's parameter count, trailing untied growth only by a constant factor. FLOPs- and depth-matched paired comparisons, with the distinction drawn in Figure 3's right panel by whether an arm moves in fitted exponent γ versus constant log A.

In data-constrained multi-epoch training the optimal loop count grows with compute, and scaling the loop count is more compute-efficient than scaling model size. Prior work shows the optimal weight decay increases with parameter count on repeated data and must be retuned at every scale; here, training ten epochs on 100M unique FineWeb tokens, the optimal loop count rises from about one to about seven, and scaling loops at fixed model size and fixed weight decay matches the best tuned model-size-scaling loss with 2.2× less compute, with weight-decay tuning on top gaining at most 0.02. A grid from Operator-1 through Loop-12 at 120M–1.4B parameters with a weight-decay sweep; optimal weight decay rises with depth but is nearly flat in loop count, and untied looping does not beat tuned Operator-1.

Perspective

The results are meant for single-epoch, data-unconstrained compute-optimal pretraining and for the data-constrained setting of multiple epochs over a fixed pool of 100M unique tokens; the authors recommend calibrating block allocation, loop count, and growth fraction at small scale and reusing them as model size and tokens scale proportionally, and they stress that hyperparameters must be tuned and scaled per architecture, since otherwise an exponent improvement can be hidden. It also points forward: context length, the number of experts, and width all grow with compute yet are fixed before training, so growing them on the schedule the network needs may change the exponent, and staged depth schedules (two passes, then four, then six) are named as a further possibility.

A careful reader would still watch how far the exponent improvement persists at larger compute and more data, for which the authors offer a single held-out run at 8× the fitted compute and a projection to 10^25 FLOPs; the roughly 20× compute gap to GPT-3 13B is described by the authors as indicative rather than a controlled comparison because the models were trained on different data and the GPT-3 reference is an estimate from a separate evaluation pipeline; KL effective depth is only a proxy for computational depth, since falling within the threshold does not mean later blocks stop changing the prediction and unused blocks in the middle of the network go undetected; random recurrence chiefly reduces the penalty for extra test-time passes rather than providing sustained test-time scaling; and the difference between Muon and Adam under width-only versus coupled width/depth scaling is flagged by the authors as a direction needing further investigation.

Sources