Public articles linked to the same research event.
arXiv For Muon under momentum accumulation and a tuned finite-step Newton–Schulz update, the work proves that full-batch training of a sufficiently wide two-layer ReLU network reaches any target empirical squared loss ε with high probability when the constant learning rate scales as (1-μ)√ε, with a hitting-time bound of O((1-μ)^{-1}ε^{-1/2}) and a sufficient width independent of both target accuracy and momentum.
For Muon under momentum accumulation and a tuned finite-step Newton–Schulz update, the work proves that full-batch training of a sufficiently wide two-layer ReLU network reaches any target empirical squared loss ε with high probability when the constant learning rate scales as (1-μ)√ε, with a hitting-time bound of O((1-μ)^{-1}ε^{-1/2}) and a sufficient width independent of both target accuracy and momentum.
For Muon under momentum accumulation and a tuned finite-step Newton–Schulz update, the work proves that full-batch training of a sufficiently wide two-layer ReLU network reaches any target empirical squared loss ε with high probability when the constant learning rate scales as (1-μ)√ε, with a hitting-time bound of O((1-μ)^{-1}ε^{-1/2}) and a sufficient width independent of both target accuracy and momentum.
For Muon under momentum accumulation and a tuned finite-step Newton–Schulz update, the work proves that full-batch training of a sufficiently wide two-layer ReLU network reaches any target empirical squared loss ε with high probability when the constant learning rate scales as (1-μ)√ε, with a hitting-time bound of O((1-μ)^{-1}ε^{-1/2}) and a sufficient width independent of both target accuracy and momentum.