Muon's Spectral Orthogonalization Lets a Linear Transformer Learn Facts Faster, Cutting the Learning-Time Ratio from √(S/R) to Constant Order
Related research and updatesSynopsis
This work studies the mechanism of Muon's spectral orthogonalization with a tractable factual-recall model: with S subjects and R relations and S>R, gradient flow learns relations before subjects with a learning-time ratio of Θ̃(√(S/R)), while spectral gradient flow reduces that ratio to Θ̃(1) and makes subject and relation errors decay as 1/(T log T) and exp(-poly(T)) respectively; gradient flow and spectral gradient flow are equivariant under orthogonal transformations of token embeddings whereas Sign gradient flow is not, so different orthonormal embeddings can yield no feature separation, a large separation phase, or even a reversed learning order.
Figure 8: SGD, AdamW, and Muon in realistic factual recall task learned by Pythia 35M. The solid line stands for the overall error rate, the dashed line stands for the error probability of the prediction having the wrong subject entity, and the dotted line stands for the error probability of the prediction having the wrong relation entity.
arXivInterpretation
In the factual-recall model, spectral orthogonalization sharply compresses the gap between learning subject information and relation information: gradient flow has a learning-time ratio of Θ̃(√(S/R)), while spectral gradient flow lowers it to Θ̃(1). Prior work (Nichani et al., 2025) showed that when S>R gradient flow learns relation-dependent information before subject-dependent information, producing a feature-separation phase, but did not characterize the timescale of that separation under Muon-like updates; this work quantifies the separation as the ratio of learning times at which the subject and relation components reach a target accuracy. Based on continuous-time limit analysis of a linear transformer on a factual-recall task, giving asymptotic learning-time ratios under gradient flow, spectral gradient flow, and Sign gradient flow; this is a theoretical characterization rather than a large-scale empirical measurement.
For fixed S and R, subject and relation errors decay as 1/(T log T) in training time T under gradient flow, but as exp(-poly(T)) under spectral gradient flow. Moves the comparison between the two optimizers from final performance to the rate at which error decays with training time, indicating that spectral orthogonalization changes the convergence order rather than only a constant factor. Derived analytically from continuous-time gradient flow and spectral gradient flow, giving functional forms for error decay; the conclusion applies within this factual-recall model setting.
Gradient flow and spectral gradient flow are equivariant under orthogonal transformations of token embeddings, whereas Sign gradient flow is not; different orthonormal embeddings can produce no feature separation, a large feature-separation phase, or even a reversed learning order. Reveals that Adam-like Sign updates lack equivariance under orthogonal embedding transformations, so learning dynamics depend on the specific orthonormal embedding chosen, while Muon-like spectral updates do not. Established through equivariance analysis and constructive argument, constituting a theoretical distinction among the three update rules.
Perspective
The results are aimed at readers studying optimizer mechanisms and feature-learning dynamics, especially those interested in the difference between Muon and Adam-like updates. The conclusions hold in the factual-recall setting: a fact maps each subject-relation pair to an answer, and a linear transformer learns the subject- and relation-dependent information needed to recover that mapping, with optimization described by three continuous-time limits—gradient flow, spectral gradient flow, and Sign gradient flow. Within this scope it gives quantitative predictions for learning-time ratios and error decay, and shows that spectral orthogonalization can make learning dynamics invariant to orthogonal transformations of token embeddings.
Readers should keep in mind that these conclusions rest on continuous-time limits and a simplified linear-transformer setting, so the gap to discrete optimization, nonlinear networks, and real pretraining remains to be clarified; the practical effect of constants and logarithmic factors in the learning-time ratio and error decay at finite scale needs further observation; and how general the reversed learning order under different orthonormal embeddings is in more complex tasks remains an open question. This summary is based on the abstract only; theorem conditions, proof details, and any experimental figures in the full text are not included, so the complete boundary of applicability still requires consulting the original.
