Skip to main content
Back to timeline
arXivSource publication:

MuonIO unifies embedding-table and language-model-head updates via 2→∞ and 1→2 operator norms, halving optimizer-state memory and cutting update FLOPs by about 46% while improving validation perplexity in 1B LLaMA pretraining

Related research and updates

Synopsis

MuonIO targets the embedding table E and language-model head L that standard Muon excludes, motivating a 1→2 operator norm for E and a 2→∞ operator norm for L, and using the identity ‖L‖_{2→∞}=‖Lᵀ‖_{1→2} to place both matrices in one vocabulary-oriented geometry under a single normalized-vector rule (column normalization for E, row normalization for L); in 1B LLaMA pretraining on C4 it reduces I/O optimizer-state memory by 50% and I/O update FLOPs by about 46% versus Muon while improving validation perplexity.

Source-provided article image: MuonIO: Principled Norm-Aware Descent for Embedding Tables and Language Model Heads
Figure 1 ·

Figure 1: Validation loss trajectories on C4 at the 130M scale for Plain Muon (baseline 1), Muon Tuned (baseline 2), Ember and MuonIO. All four configurations use standard Muon hidden updates. The inset highlights the 0.7–0.9B training token range.

arXiv

Interpretation

MuonIO provides a single Muon-style update rule for the two layer types that standard Muon excludes: the embedding table and the language-model head. Standard Muon implementations hand the input embedding table and output language-model head to AdamW, whereas MuonIO brings both layers into the same operator-norm-based principled treatment. The abstract presents a method-level derivation and unification, together with empirical evaluation in 1B LLaMA pretraining on C4.

The language-model head L uses the 2→∞ operator norm, motivated by the Lipschitz continuity of the softmax output geometry, while the embedding table E uses the 1→2 operator norm, motivated by the one-hot input geometry identified by Bernstein & Newhouse (2025). Each layer's geometry is mapped to a distinct operator norm rather than reusing the spectral-norm argument for dense linear layers. The abstract grounds the motivation in geometric arguments and a cited reference, with empirical evaluation supporting the effect.

The identity ‖L‖_{2→∞}=‖Lᵀ‖_{1→2} puts both matrices in the same vocabulary-oriented geometry, so MuonIO needs only one normalized-vector rule: column normalization for E and row normalization for L. Two treatments that previously sat on the input and output sides are merged into a single rule, simplifying the update form at the implementation level. The unification follows mathematically from the operator-norm identity; the abstract does not expand the proof details.

In 1B LLaMA pretraining on C4, MuonIO reduces I/O optimizer-state memory by 50% and I/O update FLOPs by about 46% versus Muon while improving validation perplexity. It lowers optimizer-state memory and update computation while maintaining or improving validation perplexity, pointing to resource efficiency in large-scale pretraining. The abstract reports empirical results for a single setting (1B LLaMA, C4 pretraining) without comparisons across multiple scales or datasets.

Perspective

The work targets practitioners pretraining Transformer language models with Muon-style optimizers, especially where optimizer-state memory and update computation must be controlled. The reported results apply to the 1B LLaMA pretraining setting on C4; the method itself starts from operator norms and geometric arguments, so its stated scope is the embedding table and the language-model head. For readers who want to extend Muon's principled treatment to input and output layers, MuonIO offers a directly comparable single normalization rule.

The abstract does not give the specific validation-perplexity values, training steps, batch size, or hardware configuration, nor does it state how the 50% memory and about 46% FLOPs figures were measured. The proof details of the identity ‖L‖_{2→∞}=‖Lᵀ‖_{1→2}, the concrete implementation of column and row normalization, and behavior at larger scales or other architectures all require the full text. Because this assessment is based on the abstract only, those details cannot be confirmed from the available text.

Sources