Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Fisher-guided submodular data selection lets 1B selected tokens beat 10B replay tokens on both adaptation and forgetting in medical continual pre-training

The work shows that loss-based data selection drives the post-continual-pre-training Fisher diagonal downward on exactly the high-Fisher coordinates the pretrained model had committed to while leaving low-Fisher coordinates largely untouched, and proposes a Fisher-aware selector that decomposes each candidate's gradient into an anchor component and a frontier component and aggregates them with a log-determinant submodular objective optimized in a single streaming pass; on TinyLlama-1.1B and Llama-3.1-8B continual pre-training over medical data it improves target-domain quality while bounding forgetting on held-out pretraining benchmarks, and 1B selected tokens outperform a replay strategy trained with 10B tokens on both adaptation and forgetting, a roughly 10x token-efficiency advantage.