Public articles linked to the same research event.
arXiv The work shows that loss-based data selection drives the post-continual-pre-training Fisher diagonal downward on exactly the high-Fisher coordinates the pretrained model had committed to while leaving low-Fisher coordinates largely untouched, and proposes a Fisher-aware selector that decomposes each candidate's gradient into an anchor component and a frontier component and aggregates them with a log-determinant submodular objective optimized in a single streaming pass; on TinyLlama-1.1B and Llama-3.1-8B continual pre-training over medical data it improves target-domain quality while bounding forgetting on held-out pretraining benchmarks, and 1B selected tokens outperform a replay strategy trained with 10B tokens on both adaptation and forgetting, a roughly 10x token-efficiency advantage.
The work shows that loss-based data selection drives the post-continual-pre-training Fisher diagonal downward on exactly the high-Fisher coordinates the pretrained model had committed to while leaving low-Fisher coordinates largely untouched, and proposes a Fisher-aware selector that decomposes each candidate's gradient into an anchor component and a frontier component and aggregates them with a log-determinant submodular objective optimized in a single streaming pass; on TinyLlama-1.1B and Llama-3.1-8B continual pre-training over medical data it improves target-domain quality while bounding forgetting on held-out pretraining benchmarks, and 1B selected tokens outperform a replay strategy trained with 10B tokens on both adaptation and forgetting, a roughly 10x token-efficiency advantage.
The work shows that loss-based data selection drives the post-continual-pre-training Fisher diagonal downward on exactly the high-Fisher coordinates the pretrained model had committed to while leaving low-Fisher coordinates largely untouched, and proposes a Fisher-aware selector that decomposes each candidate's gradient into an anchor component and a frontier component and aggregates them with a log-determinant submodular objective optimized in a single streaming pass; on TinyLlama-1.1B and Llama-3.1-8B continual pre-training over medical data it improves target-domain quality while bounding forgetting on held-out pretraining benchmarks, and 1B selected tokens outperform a replay strategy trained with 10B tokens on both adaptation and forgetting, a roughly 10x token-efficiency advantage.
The work shows that loss-based data selection drives the post-continual-pre-training Fisher diagonal downward on exactly the high-Fisher coordinates the pretrained model had committed to while leaving low-Fisher coordinates largely untouched, and proposes a Fisher-aware selector that decomposes each candidate's gradient into an anchor component and a frontier component and aggregates them with a log-determinant submodular objective optimized in a single streaming pass; on TinyLlama-1.1B and Llama-3.1-8B continual pre-training over medical data it improves target-domain quality while bounding forgetting on held-out pretraining benchmarks, and 1B selected tokens outperform a replay strategy trained with 10B tokens on both adaptation and forgetting, a roughly 10x token-efficiency advantage.