TACO optimizer cuts LLM fine-tuning optimizer-state memory 174x and fine-tunes 30-32B models on a single 80 GB H100
Related research and updatesSynopsis
The authors propose TACO, an optimizer that computes the exact steepest-descent direction under a dimension-normalized 1-to-1 operator norm by taking the sign of the largest-magnitude entry in each column of two-dimensional weight matrices, keeping only a small set of low-precision gradient components per column; on OPT-13B it reduces persistent optimizer state 174x versus AdamW8bit (27.7 GB to 0.16 GB) and peak training memory 2.9x (80.6 GB to 27.5 GB) with comparable accuracy and runtime, and it enables full-parameter fine-tuning of 30-32B-parameter models on a single 80 GB H100.
Figure 1: Performance–memory overview of TACO and other optimizers. Left: Pareto frontier memory-accuracy trade-off for fine-tuning OPT-13B on the SST-2 dataset on an 80 GB H100 GPU. TACO lies at the low-memory end of the Pareto frontier, achieving 94.22% accuracy with 27.54 GB of peak training memory. The dashed line connects the Pareto-optimal methods. The vertical line marks the 80 GB H100 memory limit. Adam achieves 95.3% accuracy with 254.8 GB of peak memory. Right: Peak GPU memory decomposition of TACO on OPT-30B across 8 downstream tasks. The bars show model state, optimizer state, and other transient memory across all eight downstream tasks. TACO’s optimizer state occupies only 0.267 GB, less than 0.5% of peak memory on each task, while the transient memory accounts for the variation across tasks.
arXivInterpretation
TACO computes the exact steepest-descent direction under a dimension-normalized 1-to-1 operator norm by selecting the sign of the largest-magnitude entry in each column of two-dimensional weight matrices. Muon reduces optimizer memory through matrix-valued updates, but its geometry differs from AdamW and can degrade performance when fine-tuning AdamW-pretrained models; TACO follows Muon's operator-norm steepest-descent view while taking the geometric route further to a column-wise one-sparse form, and it retains first-order gradients. The abstract states the construction and that it retains first-order gradients while making optimizer state memory nearly negligible; derivation details and convergence proofs are not provided in the abstract.
The practical TACO optimizer maintains only a small set of low-precision gradient components per column, reducing persistent optimizer state 174x versus AdamW8bit (27.7 GB to 0.16 GB) and peak training memory 2.9x (80.6 GB to 27.5 GB) on OPT-13B, with comparable accuracy and runtime. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state; TACO compresses state to nearly negligible levels while retaining first-order gradients. The abstract reports specific memory figures on OPT-13B and a comparable-accuracy-and-runtime conclusion, but does not list tasks, metrics, or statistical details.
TACO enables full-parameter fine-tuning of 30-32B-parameter models on a single 80 GB H100 across multiple model families and tasks. Optimizer-state memory overhead previously limited the model sizes that fit on modern GPUs for full-parameter fine-tuning; TACO pushes that threshold to the 30-32B range. The abstract summarizes this as 'multiple model families and tasks' without naming specific models, tasks, or evaluation scores.
Perspective
The result targets researchers and engineers who need full-parameter LLM fine-tuning under limited GPU memory, in settings where optimizer state for two-dimensional weight matrices is the constraint; the abstract's evidence comes from OPT-13B memory and accuracy comparisons and from 30-32B models fine-tuned on a single 80 GB H100 across multiple model families and tasks. Work building on its operator-norm steepest-descent view could explore how column-wise one-sparse updates behave under other optimization geometries and model scales.
The abstract does not specify the tasks, metrics, or statistical basis behind 'comparable accuracy,' nor the model families and tasks used for the 30-32B fine-tuning. How column-wise one-sparse updates behave outside the settings described, and how low-precision gradient components affect training stability, remain open questions for readers. Because this assessment is based on the abstract only, figures and experimental details in the full text could not be checked.
