Interpolation-localized middle layers plus weight freezing let gemma-3-4b multilingual CPT beat the base model on average for reading comprehension
Related research and updatesSynopsis
Using gemma-3-4b continual pretraining (CPT) across five language families, the authors interpolate pre- and post-CPT model states to localize forgetting, finding that middle-layer reversion gives the largest comprehension recovery while translation effects vary by family and direction; guided by this, they compare layer freezing, layer-range L2 regularization, post-hoc layer reversion, and model souping, and find that preserving the interpolation-identified layer weights substantially reduces comprehension loss relative to joint CPT, with layer freezing exceeding base model performance on average, while translation results are mixed.
Figure 1: (a) CPT baselines and targeted parameter alignment methods. (b) Held-in changes relative to Base. D-Rev and E-Rev restore Base weights after CPT; Freeze holds the selected window G G fixed during CPT, and L2-SP regularizes it. The targeted methods use G = [ 10 , 16 ) G=[10,16) for Belebele and G = [ 5 , 11 ) G=[5,11) for FLORES. The FLORES panels show translation directions separately and average the named families within each group.
arXivInterpretation
The authors use interpolation between pre- and post-CPT model states to localize forgetting, finding that middle-layer reversion yields the largest reading-comprehension recovery, while translation effects vary by language family and direction. Moves the question of which layers drive catastrophic forgetting in multilingual adaptation from whole-model performance comparison to layer-level localization, and offers a reusable interpolation-based diagnostic. Evaluated on gemma-3-4b across five language families, with reading comprehension and translation as the tasks, comparing reversion at different layers via pre/post-state interpolation.
When CPT strategies are guided by this layer information, preserving the identified layer weights substantially reduces comprehension loss relative to joint CPT, and layer freezing exceeds base model performance on average. Turns layer localization into training-time interventions (layer freezing, layer-range L2 regularization) and post-hoc interventions (layer reversion, model souping), benchmarked against joint multilingual and family-specific vanilla CPT baselines. Four strategy families compared against two baselines within one framework, reporting average differences on comprehension tasks.
On translation these strategies give mixed results, with dense training or post-hoc reversion often outperforming both training-time constraints and family-specific specialization. Shows that comprehension and translation respond differently to the same intervention, complicating prior assumptions that family-specific specialization alone aligns models well when extended to new tasks. Translation results differ by language family and direction, with no single strategy dominating across the board.
The authors argue that multilingual adaptation strategy should be informed by target language, base model knowledge, and downstream task, and propose interpolation-based localization as a diagnostic before committing to a training-time intervention in a new setting. Shifts from a single best strategy toward a condition-dependent decision framework with a diagnose-then-intervene order of operations. Conclusions rest on the above comparison across five language families, two task types, and multiple strategies.
Perspective
The work targets multilingual adaptation by continual pretraining gemma-3-4b across five language families, with evaluation focused on reading comprehension and translation. Its conclusions apply to practitioners who need to extend a model to new languages while preserving existing capabilities, especially teams willing to run a layer-localization diagnostic before training. Interpolation-based localization is proposed as a diagnostic step before committing to a training-time intervention in a new setting, so its value lies in guiding strategy choice rather than prescribing fixed layer indices across models.
Conclusions rest on gemma-3-4b and five language families; whether layer localization is consistent on other base models, languages, and tasks remains open. Translation results vary by family and direction and are mixed, indicating that intervention effects are condition-dependent and should be re-checked on one's own target language and task. In addition, the abstract-level text does not give concrete numbers or statistical details for each strategy, so engineering decisions should consult the full experimental tables in the original.
