Skip to main content
Back to timeline
arXivSource publication:

Nonlinear attention makes loss decay inversely with depth, while linear attention stays bound to the data spectrum

Synopsis

In controlled in-context learning tasks comparing linear and nonlinear attention, the authors find that nonlinear attention can selectively focus on relevant tokens, letting strong and weak spectral directions be learned in parallel and producing inverse-depth loss decay, whereas linear attention's depth scaling depends on the data spectrum and quickly plateaus.

Source-provided article image: Emergent Inverse-Depth Scaling From Nonlinearity In Attention
Figure 1 ·

Figure 1: The nonlinear-attention model can yield inverse-depth scaling across all tested data spectra. This overview compares depth-dependent performance across models and tasks. Colors distinguish data spectra, and all panels are plotted on a log-log scale.

arXiv

Interpretation

On the difficult top-k aggregation task, the nonlinear-attention model's evaluation loss keeps decreasing with depth with fitted exponents close to one, i.e. inverse-depth scaling, while the linear-attention model plateaus within one or two layers at a spectrum-dependent level. Prior linear-attention theory tied depth scaling to a power-law data spectrum, and empirical work reported inverse-depth scaling without explaining its origin; this work attributes inverse-depth scaling to attention nonlinearity. Controlled Transformers trained to convergence across independently varied input and teacher spectra, with log-space least-squares fits of loss versus depth whose exponents are close to one across tested spectra.

The loss admits an exact decomposition into a cross-layer shared-error term and a layer-disagreement term: the shared error converges to a spectrum-dependent constant that sets the irreducible plateau, while the disagreement term is reduced by averaging by a factor of depth, yielding continued gains. The decomposition links inverse-depth improvement to central-limit-theorem-style ensemble averaging and explains why layers behave similarly. An exact identity derived under the controlled model's parallel representation, with both terms tracked across data spectra; the shared-error term fitted separately also decays with exponents close to one.

The nonlinear-attention model reaches a substantially lower irreducible loss floor than the linear-attention model on the difficult task, e.g. roughly 2.007 to 0.341 versus 59.853 to 2.120 across the listed spectral settings. Quantifies the gap in achievable minimum loss between the two attention types, showing that selective focusing directly lowers irreducible error. Table 1 reports loss-floor values for both models across several spectral settings.

After relaxing the two-stream encoding, shared value matrix, and fixed-context assumptions toward standard Transformers, the generalized models still show approximate inverse-depth scaling signals. Extends the controlled analysis to shared representation spaces, layer-decoupled value matrices, and context updates. Appendix C relaxes each assumption in turn and repeats the experiments, observing shared-error and disagreement terms approaching finite constants and an approximate inverse-depth loss trend, though the authors note the accessible depth ranges are limited and treat these as evidence of similar scaling signals.

Perspective

The result is aimed at readers studying the mechanisms of depth scaling, and applies to controlled in-context learning settings, especially induction-head-style retrieval and top-k aggregation tasks. It suggests that when attention is nonlinear and can selectively focus on a few relevant tokens, depth scaling may depend less on the global covariance structure of the data, which is especially relevant for sparse, sharply concentrated target distributions. The authors note that similar signals appear after relaxing the two-stream encoding, shared value matrix, and fixed-context assumptions, so the mechanism may also be relevant to practical models with nonlinear attention.

The authors state that the simplified architecture omits components of full LLMs, so the identified mechanism is not guaranteed to operate there; architectural extensions show similar scaling signals but lack a clear theoretical account; experiments are restricted to in-context learning tasks, leaving generalization to real-world datasets and broader task classes untested; the mechanism need not be universal or dominant, and other processes may produce different scaling or interact with it; the analysis also does not address how depth scaling couples to width and dataset size. In addition, top-k exponents deviate more from one at the depths considered, so the asymptotic inverse-depth regime may only appear at greater depths, which remains an open question.

Sources