Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

MuonIO unifies embedding-table and language-model-head updates via 2→∞ and 1→2 operator norms, halving optimizer-state memory and cutting update FLOPs by about 46% while improving validation perplexity in 1B LLaMA pretraining

MuonIO targets the embedding table E and language-model head L that standard Muon excludes, motivating a 1→2 operator norm for E and a 2→∞ operator norm for L, and using the identity ‖L‖_{2→∞}=‖Lᵀ‖_{1→2} to place both matrices in one vocabulary-oriented geometry under a single normalized-vector rule (column normalization for E, row normalization for L); in 1B LLaMA pretraining on C4 it reduces I/O optimizer-state memory by 50% and I/O update FLOPs by about 46% versus Muon while improving validation perplexity.