Task-vector variance predicts merge collapse, and PRISM repairs five destructive merges without data
Related research and updatesSynopsis
The work proposes using the variance of specialists' task vectors across specialists as a measure of interference, yielding a pre-merge score that predicts collapse, and designs PRISM: an operator that averages task vectors first and then soft-thresholds each layer at a level set by that layer's interference; across twenty-two merge configurations from four model families only destructive merges exceed the score threshold, twelve of fourteen pre-evaluation predictions were correct, and PRISM keeps all five destructive merges above the threshold within evaluation noise of the base model without data or tuning, where plain averaging falls at least 14.4 points below it or collapses entirely.
Figure 2: Coefficient sweep on Qwen2.5-7B Math + + Coder: (a) six-task evaluation and (b) separately prompted generation. Squares mark the norm-matched control.
arXivInterpretation
The authors show that the power averaging removes equals the variance of the task vectors across specialists, which they define as interference, and derive a pre-merge score: under a working noise model, the disturbance a merge injects grows with the merge coefficient and with interference. Common merge operators give no warning before evaluation, whereas this turns collapse risk into a statistic computable before merging. Across twenty-two merge configurations from four model families, only destructive merges exceed the threshold on this score.
The authors find that statistics of sign conflict between specialists, a common target of existing merge operators, are anti-predictive. This indicates that targeting sign conflict points in a different direction from collapse risk. The text describes the observation as anti-predictive without reporting specific values.
The authors predicted the outcomes of fourteen merges before evaluating them, and twelve predictions were correct, including the destructive outcome of a specialist pair pushed past the threshold by continued pretraining. This moves the score from post-hoc explanation to pre-evaluation prediction and covers the case where continued pretraining changes a merge's outcome. Twelve of fourteen merges were predicted correctly, a prospective prediction check.
The authors introduce PRISM, which averages task vectors first and then soft-thresholds each layer at a level set by the layer's interference; without data or tuning it keeps all five destructive merges above the threshold within evaluation noise of the base model, where plain averaging falls at least 14.4 points below it or collapses entirely. The repair is calibrated by the same interference statistic, and PRISM is applied only above the threshold while merges below it, including all fifteen harmless ones, keep the plain average. All five destructive merges were repaired, with gaps to the base model within evaluation noise.
Perspective
The results target specialists fine-tuned from a shared base and merged by averaging task vectors, applying to pre-merge risk assessment and to switching to PRISM when the threshold is exceeded; merges below the threshold, including all fifteen harmless ones, keep the plain average. Code is released, supporting reproduction and adoption in comparable merging pipelines.
The abstract does not give the functional form linking interference and the merge coefficient, how the threshold is set, or the composition of the twenty-two configurations and four model families and the evaluation metric; the two incorrect cases among the fourteen pre-evaluation predictions are not detailed. Implementation would still require the paper's noise-model derivation, threshold setting, and per-layer soft-thresholding details.
