Public articles linked to the same research event.
arXiv MuonIO targets the embedding table E and language-model head L that standard Muon excludes, motivating a 1→2 operator norm for E and a 2→∞ operator norm for L, and using the identity ‖L‖_{2→∞}=‖Lᵀ‖_{1→2} to place both matrices in one vocabulary-oriented geometry under a single normalized-vector rule (column normalization for E, row normalization for L); in 1B LLaMA pretraining on C4 it reduces I/O optimizer-state memory by 50% and I/O update FLOPs by about 46% versus Muon while improving validation perplexity.
MuonIO targets the embedding table E and language-model head L that standard Muon excludes, motivating a 1→2 operator norm for E and a 2→∞ operator norm for L, and using the identity ‖L‖_{2→∞}=‖Lᵀ‖_{1→2} to place both matrices in one vocabulary-oriented geometry under a single normalized-vector rule (column normalization for E, row normalization for L); in 1B LLaMA pretraining on C4 it reduces I/O optimizer-state memory by 50% and I/O update FLOPs by about 46% versus Muon while improving validation perplexity.
MuonIO targets the embedding table E and language-model head L that standard Muon excludes, motivating a 1→2 operator norm for E and a 2→∞ operator norm for L, and using the identity ‖L‖_{2→∞}=‖Lᵀ‖_{1→2} to place both matrices in one vocabulary-oriented geometry under a single normalized-vector rule (column normalization for E, row normalization for L); in 1B LLaMA pretraining on C4 it reduces I/O optimizer-state memory by 50% and I/O update FLOPs by about 46% versus Muon while improving validation perplexity.
MuonIO targets the embedding table E and language-model head L that standard Muon excludes, motivating a 1→2 operator norm for E and a 2→∞ operator norm for L, and using the identity ‖L‖_{2→∞}=‖Lᵀ‖_{1→2} to place both matrices in one vocabulary-oriented geometry under a single normalized-vector rule (column normalization for E, row normalization for L); in 1B LLaMA pretraining on C4 it reduces I/O optimizer-state memory by 50% and I/O update FLOPs by about 46% versus Muon while improving validation perplexity.