TESS replaces per-sample weights with a Pointwise Value Matching objective, letting a data-selection network transfer across datasets and model scales in LLM safety and instruction tuning
Related research and updatesSynopsis
The work reports that directly incorporating a selection network into existing meta-learning for training-data selection (MTS) objectives causes unstable optimization and poor generalization due to weight suppression and persistent reliance on easy-to-learn features, and proposes Transferable Example Scoring and Selection (TESS), a scalable data-selection framework built on a Pointwise Value Matching (PVM) objective, which in experiments on LLM safety and targeted instruction tuning demonstrates strong transfer across datasets, from subsets to full corpora, and from smaller to larger models.
Figure 1: Simulation of the weight dynamics in Theorem 1 . Curves show the weights w i ( t ) w_{i}(t) , with fixed pseudo-labels Δ i ∗ \Delta_{i}^{*} indicated in the legend. Initial weights are ( 0.2 , 0.2 , 0.3 , 0.2 ) (0.2,0.2,0.3,0.2) , and δ 0 \delta_{0} is 10 − 6 10^{-6} .
arXivInterpretation
Directly using a selection network with existing MTS objectives leads to unstable optimization and poor generalization. The authors attribute this to weight suppression and persistent reliance on easy-to-learn features, identifying the combination of a selection network with existing MTS objectives as a case that requires a different loss. The abstract states this as 'we find', an observation from the authors' own experiments, without reporting specific metrics or ablation sizes in the abstract.
It proposes TESS, a scalable data-selection framework built on a Pointwise Value Matching (PVM) objective. Unlike per-sample weights or directly embedding a selection network, TESS uses PVM as the learning objective to support transferable example scoring and selection. The abstract provides the framework name, the objective name, and the design motivation, but does not detail the mathematical form of PVM or the training procedure.
On LLM safety and targeted instruction tuning, TESS shows strong transfer across datasets, from subsets to full corpora, and from smaller to larger models. The transfer dimensions cover both data distribution and model scale, addressing the trade-off in existing methods between fine-grained valuation and transferability to unseen data. The abstract summarizes the experimental conclusion as 'demonstrate strong transfer' without listing dataset names, model sizes, baselines, or numerical values.
Perspective
The work targets research and engineering teams that train LLMs on massive heterogeneous corpora, especially the data-selection step for safety alignment and targeted instruction tuning. The results described in the abstract apply to its tested safety and instruction-tuning settings and are claimed to transfer from subsets to full corpora and from smaller to larger models; readers can therefore treat it as a candidate scalable, transferable data-selection framework to replace per-sample weights or heuristic scoring.
The abstract does not specify the mathematical form of PVM, the training and inference cost of the selection network, how weight suppression is diagnosed, or dataset names, model sizes, baseline comparisons, and numerical results. Readers should therefore still watch whether TESS transfers beyond the safety and instruction-tuning tasks described, over what scale range the smaller-to-larger model transfer holds, and what makes the PVM objective more stable than existing MTS objectives. Because this assessment is based only on the abstract, figures and experimental details could not be checked, and these questions remain open.
