Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

LSL gates weight updates to the current data's support with a GMM, letting a 7B-parameter LLM keep 96.6% of pretrained capability after multi-phase finetuning

The authors propose Local Support Learning (LSL), which frames catastrophic forgetting as a geometric problem in each weight matrix's input space and uses a Gaussian Mixture Model gate to enable a weight adapter only on the support of the current phase's training distribution, so that LLMs of up to 7 billion parameters can be finetuned across multiple phases without access to prior data while learning new tasks at full capacity and retaining pretrained and previously finetuned capabilities.