SVC selects singular-vector channels to raise continual-learning performance while preserving pretrained ability across four LLM families
Synopsis
The work defines each singular value's paired left and right singular vectors as a singular-vector channel, estimates each channel's adaptation benefit from current-task gradients and its forgetting cost from activation second moments on an unlabeled general-domain corpus, then selects trainable channels via knee-based cost screening, Pareto-front filtering, and Otsu thresholding; across four LLM families, eight TRACE tasks, and three task orders, SVC preserves pretrained general capabilities while achieving strong downstream continual-learning performance.
Figure 1: Stability–plasticity trade-off between LoRA and SVF. (a) GSM8K retention across fine-tuning tasks. (b) Downstream performance. Details in Appendix C
arXivInterpretation
The paper proposes the singular-vector channel as the basic unit for controlling the stability-plasticity trade-off: each channel is a paired left and right singular-vector direction that can be updated to acquire new knowledge or fixed to preserve pretrained capabilities. Prior PEFT methods such as LoRA, O-LoRA, SVF, SVFT, PiSSA, and MiLoRA specify the set of trainable spectral coefficients before fine-tuning, or constrain orthogonality only relative to previous-task low-rank subspaces, rather than using task-specific estimates of adaptation benefit and forgetting cost to decide which directions are trainable. The unit and its motivation come from a comparison of LoRA and SVF: when the same pretrained model is independently adapted to each of the eight TRACE tasks, LoRA shows markedly different drops in GSM8K, whereas SVF's directional constraint yields more consistent retention but limits downstream adaptation.
Before fine-tuning, SVC estimates each channel's adaptation benefit and forgetting cost, then automatically selects a task-specific trainable channel subset through three stages: knee-based cost screening, Pareto-front filtering, and Otsu thresholding. An ideal evaluation of forgetting cost requires historical activations, which are unavailable; the paper shows historical data enter the score only through the activation second-moment matrix, so it can be approximated from an unlabeled general-domain corpus, requiring neither prior-task samples nor task-specific modules. Ablations show removing Pareto-front filtering causes the largest degradation, nearly eliminating retention of code-generation and mathematical-reasoning capabilities and sharply reducing sequential performance; removing knee-based cost screening has a smaller effect but still weakens mathematical-reasoning retention, lowers sequential performance, and increases forgetting.
Across four LLM families, eight TRACE tasks, and three task orders, SVC preserves pretrained general capabilities while achieving strong continual-learning performance. Across the 12 model and task-order settings, SVC achieves or ties the best result in 10 HumanEval settings and 11 GSM8K settings; against the strongest baseline in each setting, it achieves mean gains of 1.68 percentage points on HumanEval and 5.18 percentage points on GSM8K. SVC obtains the best OP in nearly all settings, and its BWT results indicate substantially reduced forgetting; despite storing no previous-task samples, it consistently surpasses both replay baselines in OP and remains competitive with them in BWT. It ranks first or second in 46 of the 48 combinations of model, task order, and metric.
The channel scoring and selection mechanism is validated: the cost score identifies risky update directions, and the benefit score reflects a channel's contribution to current-task adaptation. Under cost-matched conditions comparing high-benefit and low-benefit channel groups, the high-benefit groups achieve higher AP in both the lowest-cost and middle-cost regions, indicating the benefit score captures a channel's contribution to current-task adaptation rather than being driven only by cost differences. When channels are selected from the highest-, middle-, or lowest-cost regions, retention of pretrained capabilities improves consistently as the cost of the selected channels decreases; high-cost selection also affects sequential learning, suggesting excessive cost can offset adaptation gains.
Perspective
The result targets LLM practitioners who need continual domain adaptation without prior-task samples or pretraining data, in settings where linear layers admit SVD and an unlabeled general-domain reference corpus is available. The method is validated on four open-source model families (Llama3-8B, Mistral-7B-v0.3, Gemma2-9B-it, Qwen3.5-9B-Base), eight TRACE tasks, and three task orders, with a fixed reference corpus of 30,000 unlabeled examples sampled from C4 and Stack-v2-Python at a 2:1 ratio. Because channel selection is completed before fine-tuning and fixed throughout training on that task, and learned increments are merged back into the weights without persistent task-specific modules, the method can be embedded directly into continual-learning pipelines that must not grow model size.
The reference corpus consists only of C4 and Stack-v2-Python at a 2:1 ratio, so its robustness to broader capability coverage and distribution shift remains to be examined; channel selection is determined once before fine-tuning and fixed during training, without modeling interactions among channels or dynamic adjustment over training; experiments cover eight TRACE tasks and three task orders, leaving longer task sequences and a broader range of general-capability benchmarks to be characterized. In addition, the benefit and forgetting-cost derivations rely on a first-order Taylor expansion and dropping second-order perturbation terms, and the ablation on whether singular values are included in the channel shows joint selection decreases HumanEval, GSM8K, OP, and BWT, which the paper attributes in part to changed score distributions and singular values continuing to adjust during training, an explanation offered as a possible cause rather than a settled conclusion.
