Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Clean Cuts Second-Order Optimizer Memory to Linear via Nyström Sketching, Enabling 13B Pre-training on a Single 80GB GPU

The authors introduce Clean and its low-precision variant Q-Clean: randomized Nyström approximations replace SOAP's explicit left and right preconditioners, reducing optimizer-state memory from quadratic to linear in layer dimensions, while four orthogonal blocks (Main, A, B, C) reintegrate off-subspace components into the update; pre-training LLaMA-350M and 1.3B on C4, Clean reaches perplexity close to SOAP with less optimizer memory than AdamW, Q-Clean cuts optimizer memory by over 50% versus Muon, Clean reaches AdamW's final performance 26% faster in wall-clock time, and the methods uniquely enable pre-training a 13B model on a single 80GB H100.