Public articles linked to the same research event.
arXiv The authors propose Local Support Learning (LSL), which frames catastrophic forgetting as a geometric problem in each weight matrix's input space and uses a Gaussian Mixture Model gate to enable a weight adapter only on the support of the current phase's training distribution, so that LLMs of up to 7 billion parameters can be finetuned across multiple phases without access to prior data while learning new tasks at full capacity and retaining pretrained and previously finetuned capabilities.
The authors propose Local Support Learning (LSL), which frames catastrophic forgetting as a geometric problem in each weight matrix's input space and uses a Gaussian Mixture Model gate to enable a weight adapter only on the support of the current phase's training distribution, so that LLMs of up to 7 billion parameters can be finetuned across multiple phases without access to prior data while learning new tasks at full capacity and retaining pretrained and previously finetuned capabilities.
The authors propose Local Support Learning (LSL), which frames catastrophic forgetting as a geometric problem in each weight matrix's input space and uses a Gaussian Mixture Model gate to enable a weight adapter only on the support of the current phase's training distribution, so that LLMs of up to 7 billion parameters can be finetuned across multiple phases without access to prior data while learning new tasks at full capacity and retaining pretrained and previously finetuned capabilities.
The authors propose Local Support Learning (LSL), which frames catastrophic forgetting as a geometric problem in each weight matrix's input space and uses a Gaussian Mixture Model gate to enable a weight adapter only on the support of the current phase's training distribution, so that LLMs of up to 7 billion parameters can be finetuned across multiple phases without access to prior data while learning new tasks at full capacity and retaining pretrained and previously finetuned capabilities.