Raising Adam's second moment only for the output projection removes 39.4%–67.9% of catastrophic forgetting across eight settings from 160M to 12B
Synopsis
In data-free continual pre-training and fine-tuning, the authors use training-time parameter freezing to localize catastrophic forgetting to the output embedding rows of tokens rare in the new corpus, show that these rows receive persistent one-sided softmax gradients that Adam's second-moment normalization amplifies into full-sized updates, and raise Adam's second moment only for the output projection, removing 39.4%–67.9% of forgetting across seven stable settings from 160M to 12B and four model families without degrading target learning.
Figure 1: Overview: where the forgetting lives, why it lives there, and what acting on it buys.
arXivInterpretation
Forgetting is localizable in parameter space: freezing low-second-moment rows of the output embedding removes 31.3%–51.0% of forgetting across five settings, whereas freezing input embeddings in the same band is negligible and the low-second-moment body band contributes only 0.4%–12.3%. Prior localization reports conflicted across early attention, lower layers, and FFNs, largely from correlational metrics like weight displacement; training-time freezing tests necessity directly and attributes causal share to specific output embedding rows. A per-group freezing grid at six settings up to 1.4B with mass-matched body controls; on Qwen, random output rows remove 12.1% at a learning cost while the lowest-second-moment coordinates remove 42.7% for free.
The site is governed by the corpus's vocabulary deficiency rather than the training mode: on a fixed base model, baseline forgetting strictly follows vocabulary deficiency (0.043, 0.080, 0.657 nats for math, code, Korean at 410M), enabling pre-retraining risk ranking from token counts alone. Shifts localization from training mode to a data-side variable and supplies a training-free diagnostic computable with the base tokenizer before retraining. Predicted ranks strictly match empirical ranks at three models and four earlier (checkpoint, corpus) comparisons; absolute-magnitude prediction failed, with a Spearman rank correlation of 0.60 across ten configurations, below the pre-registered 0.7.
Mechanistically, tokens absent as targets receive only one-sided softmax gradients, and Adam's second-moment normalization turns these minute gradients into near-full-learning-rate updates; raising the second moment only for the output projection selectively brakes them because output coordinates sit two to three orders of magnitude below body coordinates. Names Adam's second moment as both amplifier and brake and yields a single-line optimizer change requiring no explicit token list; freezing under-exposed rows and per-step gating are close equivalents across five settings. Travel-distance probes over 120 steps show 48.0–189.2-fold damping in the lowest band versus 1.4–1.8 in the highest; SVD shows absent-row displacement energy concentrating in one direction (66.1%–80.8%) against 3.2%–12.9% in high-exposure controls.
The intervention removes 39.4%–67.9% of forgetting across seven stable settings of an eight-setting grid, composes additively or better with replay (79.8% on Qwen/Korean), and rescues released-head LoRA from a 23-fold forgetting surge, while post-hoc editing of drifted rows recovers under 5% of forgetting. Against subspace methods that need past data and regularizers that blindly constrain aggregate drift, this is a zero-overhead, tuning-free defense that must operate during training. 160M to 12B across four model families at three seeds (the bimodal 410M arm read at six); extending the budget to 400M tokens raises rather than decays the reduction (6.9B from 53.9% to 62.3%, 12B from 76.2% to 84.1%).
Perspective
The result targets data-free continual pre-training and fine-tuning, especially domain injection where the new corpus brings substantial new vocabulary and vocabulary deficiency is high (Korean, code). In that setting, a reader can compute, before retraining, the share of old-domain positions whose target token is under-exposed in the new corpus: a high fraction indicates output-dominated forgetting, warranting a raised second moment on the output projection; a low fraction indicates body-dominated forgetting, where freezing harms target learning and replay or learning-rate scaling is preferable. Mitigation holds across the low-rate plateau, the optimal rate is invariant to the intervention, so one standard sweep suffices and no per-model tuning is needed; when replay is available, activate the zero-cost optimizer change first, then deploy replay to curb residual forgetting.
The vocabulary-deficiency diagnostic predicts rank order only, not absolute forgetting, and cross-tokenizer calibration did not succeed (OLMo's per-token Korean loss looks close to English, but per-byte loss reveals it as the least Korean-familiar of the four families). Under in-distribution tuning the causal share moves into the body, and body-scoped damping is lossless at only one of the two measured cells, which the authors register as preliminary. Reduction is not uniform within a setting: two of six seeds at Pythia-410M on Korean recover little and forfeit target learning. Scaling comparisons are Pythia-only, the second-moment sweep covers three Pythia checkpoints, and the capability check predates the grid under different conditions (math injection, 1% replay). The adapter ablation uses one rank at two base models. The mechanism needs coordinate-wise adaptive optimizers, leaving Muon outside its scope. Attention freezing on KoAlpaca removes 42.7% of forgetting at 410M, above the 29.8% of output-scoped, and the authors state that causal shares do not partition across unnormalized modules, so whether attention holds an independent share remains open.
