Public articles linked to the same research event.
arXiv Using gemma-3-4b continual pretraining (CPT) across five language families, the authors interpolate pre- and post-CPT model states to localize forgetting, finding that middle-layer reversion gives the largest comprehension recovery while translation effects vary by family and direction; guided by this, they compare layer freezing, layer-range L2 regularization, post-hoc layer reversion, and model souping, and find that preserving the interpolation-identified layer weights substantially reduces comprehension loss relative to joint CPT, with layer freezing exceeding base model performance on average, while translation results are mixed.
Using gemma-3-4b continual pretraining (CPT) across five language families, the authors interpolate pre- and post-CPT model states to localize forgetting, finding that middle-layer reversion gives the largest comprehension recovery while translation effects vary by family and direction; guided by this, they compare layer freezing, layer-range L2 regularization, post-hoc layer reversion, and model souping, and find that preserving the interpolation-identified layer weights substantially reduces comprehension loss relative to joint CPT, with layer freezing exceeding base model performance on average, while translation results are mixed.
Using gemma-3-4b continual pretraining (CPT) across five language families, the authors interpolate pre- and post-CPT model states to localize forgetting, finding that middle-layer reversion gives the largest comprehension recovery while translation effects vary by family and direction; guided by this, they compare layer freezing, layer-range L2 regularization, post-hoc layer reversion, and model souping, and find that preserving the interpolation-identified layer weights substantially reduces comprehension loss relative to joint CPT, with layer freezing exceeding base model performance on average, while translation results are mixed.
Using gemma-3-4b continual pretraining (CPT) across five language families, the authors interpolate pre- and post-CPT model states to localize forgetting, finding that middle-layer reversion gives the largest comprehension recovery while translation effects vary by family and direction; guided by this, they compare layer freezing, layer-range L2 regularization, post-hoc layer reversion, and model souping, and find that preserving the interpolation-identified layer weights substantially reduces comprehension loss relative to joint CPT, with layer freezing exceeding base model performance on average, while translation results are mixed.