Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Muon with finite-step Newton–Schulz orthogonalization gets its first training-loss guarantee: hitting time O((1-μ)^{-1}ε^{-1/2}) to any target loss ε

For Muon under momentum accumulation and a tuned finite-step Newton–Schulz update, the work proves that full-batch training of a sufficiently wide two-layer ReLU network reaches any target empirical squared loss ε with high probability when the constant learning rate scales as (1-μ)√ε, with a hitting-time bound of O((1-μ)^{-1}ε^{-1/2}) and a sufficient width independent of both target accuracy and momentum.