Skip to main content
Back to timeline
arXivSource publication:

Muon with finite-step Newton–Schulz orthogonalization gets its first training-loss guarantee: hitting time O((1-μ)^{-1}ε^{-1/2}) to any target loss ε

Related research and updates

Synopsis

For Muon under momentum accumulation and a tuned finite-step Newton–Schulz update, the work proves that full-batch training of a sufficiently wide two-layer ReLU network reaches any target empirical squared loss ε with high probability when the constant learning rate scales as (1-μ)√ε, with a hitting-time bound of O((1-μ)^{-1}ε^{-1/2}) and a sufficient width independent of both target accuracy and momentum.

Source-provided article image: Training-Loss Guarantees for Muon with Finite-Step Newton--Schulz Orthogonalization
Figure 1 ·

Figure 1: Training loss across accuracy targets and momentum values. (a) Loss curves for five target values of ε \varepsilon at μ = 0.95 \mu=0.95 ; dots mark the first iteration reaching each target. (b) Hitting time versus ε \varepsilon , with an ε − 1 / 2 \varepsilon^{-1/2} reference line. (c) Loss curves for five momentum values at target 10 − 2 10^{-2} , using η = ( 1 − μ ) ​ η 0 \eta=(1-\mu)\eta_{0} with η 0 = 1.54 × 10 − 4 \eta_{0}=1.54\times 10^{-4} . (d) The curves in (c) plotted against ( 1 − μ ) ​ t (1-\mu)t . Panels (a,b) use a separate data draw ( λ 0 ≈ 0.34 \lambda_{0}\approx 0.34 , R 0 ≈ 26 R_{0}\approx 26 ); panels (c,d) use the saved draw ( λ 0 = 0.335 \lambda_{0}=0.335 , R 0 = 30.2 R_{0}=30.2 ).

arXiv

Interpretation

The paper provides a finite-time training-loss guarantee for Muon rather than mere stationarity: for a sufficiently wide two-layer ReLU network with fixed random output weights and a positive-definite limiting neural tangent kernel, full-batch Muon reaches any target empirical squared loss ε>0 with high probability. Earlier Muon convergence analyses either assumed exact orthogonalization or analyzed classical Newton–Schulz polynomials and guaranteed only stationarity; this work incorporates both momentum accumulation and the tuned finite-step update, moving the conclusion from stationarity to training loss. The result is stated as a theorem with explicit conditions: sufficiently wide two-layer ReLU network, fixed random output weights, positive-definite limiting neural tangent kernel, full-batch training, holding with high probability over initialization.

For every momentum parameter μ∈[0,1), a target-dependent constant learning rate proportional to (1-μ)√ε yields a hitting-time bound of O((1-μ)^{-1}ε^{-1/2}), and the sufficient width is independent of both target accuracy and momentum. This makes the quantitative relation among learning rate, momentum, and target accuracy explicit, and shows the width condition does not tighten with accuracy requirements or momentum choice. Both the rate and the width independence are stated in the theorem with other problem parameters fixed; the abstract does not give numerical values for the constant factors.

The analysis characterizes the mechanism: the tuned Newton–Schulz map preserves alignment with the momentum buffer while bounding the update's spectral norm, and control of gradient variation near initialization transfers this alignment to the current gradient, ensuring descent until the target is reached without requiring exact orthogonalization. This turns the question of what the five tuned Newton–Schulz steps preserve from an empirical observation into a provable mechanism: what is preserved is alignment with the momentum buffer, not exact orthogonality. The mechanism comes from the theoretical analysis and is supported by numerical experiments: at widths below the sufficient theoretical threshold, gradient-update alignment remains above the analytical reference.

Numerical experiments on a fixed teacher–student dataset validate the mechanism: all 30 runs across six widths and five student initializations reach the target loss while maintaining kernel positivity. The experiments reproduce the predicted alignment and convergence behavior at widths below the sufficient theoretical threshold, indicating the mechanism also appears outside the theoretical threshold. 30 runs, six widths, five initializations, fixed dataset; the abstract reports no error bars or statistics beyond this setting.

Perspective

The result applies to the setting of a sufficiently wide two-layer ReLU network, fixed random output weights, a positive-definite limiting neural tangent kernel, and full-batch training, with a target-dependent constant learning rate proportional to (1-μ)√ε. Within this scope it shows that the tuned five-step Newton–Schulz update guarantees descent without exact orthogonalization, and that the width condition does not tighten with target accuracy or momentum, which is relevant to researchers and practitioners interested in Muon's theoretical properties and the learning-rate–momentum scaling relation. Numerical experiments further test the mechanism at widths below the sufficient theoretical threshold.

The abstract does not give numerical values for the constant factors in the theorem or the precise probability form of the high-probability statement; error ranges and statistical details of the numerical experiments are not reported in the abstract. The theory assumes two-layer ReLU, fixed output weights, a positive-definite limiting kernel, and full-batch training, so extension to deeper networks, other activations and architectures, mini-batch or adaptive learning-rate settings remains open. The abstract also does not state the specific target loss values used or the size of the teacher–student dataset, details that require the full text.

Sources