Public articles linked to the same research event.
arXiv This work studies the mechanism of Muon's spectral orthogonalization with a tractable factual-recall model: with S subjects and R relations and S>R, gradient flow learns relations before subjects with a learning-time ratio of Θ̃(√(S/R)), while spectral gradient flow reduces that ratio to Θ̃(1) and makes subject and relation errors decay as 1/(T log T) and exp(-poly(T)) respectively; gradient flow and spectral gradient flow are equivariant under orthogonal transformations of token embeddings whereas Sign gradient flow is not, so different orthonormal embeddings can yield no feature separation, a large separation phase, or even a reversed learning order.
This work studies the mechanism of Muon's spectral orthogonalization with a tractable factual-recall model: with S subjects and R relations and S>R, gradient flow learns relations before subjects with a learning-time ratio of Θ̃(√(S/R)), while spectral gradient flow reduces that ratio to Θ̃(1) and makes subject and relation errors decay as 1/(T log T) and exp(-poly(T)) respectively; gradient flow and spectral gradient flow are equivariant under orthogonal transformations of token embeddings whereas Sign gradient flow is not, so different orthonormal embeddings can yield no feature separation, a large separation phase, or even a reversed learning order.
This work studies the mechanism of Muon's spectral orthogonalization with a tractable factual-recall model: with S subjects and R relations and S>R, gradient flow learns relations before subjects with a learning-time ratio of Θ̃(√(S/R)), while spectral gradient flow reduces that ratio to Θ̃(1) and makes subject and relation errors decay as 1/(T log T) and exp(-poly(T)) respectively; gradient flow and spectral gradient flow are equivariant under orthogonal transformations of token embeddings whereas Sign gradient flow is not, so different orthonormal embeddings can yield no feature separation, a large separation phase, or even a reversed learning order.
This work studies the mechanism of Muon's spectral orthogonalization with a tractable factual-recall model: with S subjects and R relations and S>R, gradient flow learns relations before subjects with a learning-time ratio of Θ̃(√(S/R)), while spectral gradient flow reduces that ratio to Θ̃(1) and makes subject and relation errors decay as 1/(T log T) and exp(-poly(T)) respectively; gradient flow and spectral gradient flow are equivariant under orthogonal transformations of token embeddings whereas Sign gradient flow is not, so different orthonormal embeddings can yield no feature separation, a large separation phase, or even a reversed learning order.