Skip to main content
Back to timeline
arXivSource publication:

Fisher-guided submodular data selection lets 1B selected tokens beat 10B replay tokens on both adaptation and forgetting in medical continual pre-training

Related research and updates

Synopsis

The work shows that loss-based data selection drives the post-continual-pre-training Fisher diagonal downward on exactly the high-Fisher coordinates the pretrained model had committed to while leaving low-Fisher coordinates largely untouched, and proposes a Fisher-aware selector that decomposes each candidate's gradient into an anchor component and a frontier component and aggregates them with a log-determinant submodular objective optimized in a single streaming pass; on TinyLlama-1.1B and Llama-3.1-8B continual pre-training over medical data it improves target-domain quality while bounding forgetting on held-out pretraining benchmarks, and 1B selected tokens outperform a replay strategy trained with 10B tokens on both adaptation and forgetting, a roughly 10x token-efficiency advantage.

Source-provided article image: Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models
Figure 1 ·

Figure 1: Parameter-agnostic CPT erodes high-Fisher directions. Diagonal Fisher entries after CPT. Left: Parameter-agnostic CPT produces substantial negative drift on high-Fisher coordinates (orange regions), indicating damage to directions where the pretrained model had strong commitments. Right: Our Fisher-guided selector mitigates this drift while concentrating acquisition (blue regions) in low-Fisher coordinates where the pretrained model is less committed.

arXiv

Interpretation

The paper reports a parameter-space mechanism for forgetting: loss-based selection causes the post-CPT Fisher diagonal to drift downward on exactly the high-Fisher coordinates the pretrained model had committed to, while low-Fisher coordinates are left largely untouched. Existing CPT practice either scores candidates with parameter-agnostic scalars such as perplexity or mitigates forgetting by spending many extra general-domain replay tokens; neither directly asks how training on a candidate will move the model parameters, whereas this work reframes forgetting as an asymmetric drift in parameter space. The evidence is a comparison of the Fisher diagonal before and after continual pre-training, which the authors summarize as an asymmetry that 'exposes a parameter-space mechanism for catastrophic forgetting'; the abstract reports the mechanism rather than specific coordinate counts or statistics.

The authors propose a Fisher-aware CPT selector that decomposes each candidate's gradient into an anchor component, measuring perturbation along committed parameter directions, and a frontier component, measuring update capacity in unconstrained low-Fisher subspaces. Relative to parameter-agnostic scalar scoring and general-domain replay, this decomposition makes the direction of parameter movement the basis for selection, placing adaptation and forgetting signals in one framework. The method description states the definition and role of the two components, a design-level contribution; the abstract does not give the detailed computation of the components or ablations.

These signals are aggregated with a log-determinant submodular objective and optimized in a single pass using a scalable streaming data selection pipeline. Compared with approaches requiring multiple iterations or extra replay budget, the combination of this objective and a streaming pipeline allows selection in one traversal, targeting the practical constraints of web-scale corpora and finite token budgets. The abstract states the objective form and the single-pass streaming optimization; it does not report selection-stage compute numbers or comparisons against alternative objectives.

In continual pre-training over medical data on TinyLlama-1.1B and Llama-3.1-8B, the selector improves target-domain quality and bounds forgetting on held-out pretraining benchmarks; 1B selected tokens outperform a replay strategy trained with 10B tokens on both adaptation and forgetting, giving about a 10x token-efficiency advantage. Relative to forgetting-aware replay, this result reduces the token volume needed for better adaptation and less forgetting by roughly an order of magnitude, indicating that selection quality can substitute for part of the replay budget. The evidence is a comparison across two model scales, a medical target domain, and held-out pretraining benchmarks, with a 1B-versus-10B token-efficiency comparison; the abstract does not list specific metric values, benchmark names, or variance.

Perspective

The result targets continual pre-training with medical data as the target domain, validated at two model scales, TinyLlama-1.1B and Llama-3.1-8B, and evaluated on target-domain quality and forgetting on held-out pretraining benchmarks. It applies to continual pre-training pipelines with limited token budgets that must trade off adaptation against preserving general capabilities, especially where data selection is meant to replace part of general-domain replay. The method depends on estimating candidate gradients and Fisher information and performs selection in a single streaming pass, so its applicability presupposes access to these gradient signals and acceptance of a one-pass selection cost.

The abstract does not give specific evaluation metric values, held-out benchmark names, quantitative statistics of the Fisher drift, or the compute cost of the selection stage, so the robustness of the 10x token-efficiency advantage still needs confirmation in the full experimental setup of the paper. The relative weighting of the anchor and frontier components, the approximation quality of the log-determinant submodular objective, and how the mechanism behaves in target domains beyond medicine and at other model scales are open questions worth watching. From the currently visible text alone, the selector's full position relative to continual-pre-training strategies other than perplexity scoring and general-domain replay cannot be determined.

Sources