Public articles linked to the same research event.
arXiv The authors introduce Clean and its low-precision variant Q-Clean: randomized Nyström approximations replace SOAP's explicit left and right preconditioners, reducing optimizer-state memory from quadratic to linear in layer dimensions, while four orthogonal blocks (Main, A, B, C) reintegrate off-subspace components into the update; pre-training LLaMA-350M and 1.3B on C4, Clean reaches perplexity close to SOAP with less optimizer memory than AdamW, Q-Clean cuts optimizer memory by over 50% versus Muon, Clean reaches AdamW's final performance 26% faster in wall-clock time, and the methods uniquely enable pre-training a 13B model on a single 80GB H100.
The authors introduce Clean and its low-precision variant Q-Clean: randomized Nyström approximations replace SOAP's explicit left and right preconditioners, reducing optimizer-state memory from quadratic to linear in layer dimensions, while four orthogonal blocks (Main, A, B, C) reintegrate off-subspace components into the update; pre-training LLaMA-350M and 1.3B on C4, Clean reaches perplexity close to SOAP with less optimizer memory than AdamW, Q-Clean cuts optimizer memory by over 50% versus Muon, Clean reaches AdamW's final performance 26% faster in wall-clock time, and the methods uniquely enable pre-training a 13B model on a single 80GB H100.
The authors introduce Clean and its low-precision variant Q-Clean: randomized Nyström approximations replace SOAP's explicit left and right preconditioners, reducing optimizer-state memory from quadratic to linear in layer dimensions, while four orthogonal blocks (Main, A, B, C) reintegrate off-subspace components into the update; pre-training LLaMA-350M and 1.3B on C4, Clean reaches perplexity close to SOAP with less optimizer memory than AdamW, Q-Clean cuts optimizer memory by over 50% versus Muon, Clean reaches AdamW's final performance 26% faster in wall-clock time, and the methods uniquely enable pre-training a 13B model on a single 80GB H100.
The authors introduce Clean and its low-precision variant Q-Clean: randomized Nyström approximations replace SOAP's explicit left and right preconditioners, reducing optimizer-state memory from quadratic to linear in layer dimensions, while four orthogonal blocks (Main, A, B, C) reintegrate off-subspace components into the update; pre-training LLaMA-350M and 1.3B on C4, Clean reaches perplexity close to SOAP with less optimizer memory than AdamW, Q-Clean cuts optimizer memory by over 50% versus Muon, Clean reaches AdamW's final performance 26% faster in wall-clock time, and the methods uniquely enable pre-training a 13B model on a single 80GB H100.