Skip to main content
Back to timeline
arXivSource publication:

Poisoned fine-tuning services can be repaired at two layers: harmful score falls from 11.54 to 0.08 with task accuracy intact

Lead

Malicious fine-tuning erodes a model's refusal behavior, and SLDR trains a LoRA recovery adapter on only the two layers at the ends of a signed sensitivity spectrum and activates it per query, cutting the average harmful score on Llama3.1/SST2 from 11.54 to 0.08 while downstream accuracy holds.

Source-provided article image: SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing
Figure 1 · arXiv

Story

Fine-tuning-as-a-service lets users adapt an aligned model with their own data, and an attacker can use the same interface to mix a few harmful instruction-response pairs into otherwise benign task data, so the model keeps working on legitimate inputs but stops refusing harmful requests. Earlier repairs either constrained fine-tuning before harmful samples were learned or applied coarse model-level recovery afterward, such as perturbing weights or pruning harmful parameters, which entangles safety repair with task behavior and makes it hard to restore refusal without hurting task performance. In the default Llama3.1/SST2 setting, undefended SFT reaches an average harmful score of 11.54, while SLDR brings it to 0.08 with downstream accuracy matching SFT.

The method first runs a signed sensitivity diagnosis over the layers, then repairs only the two layers at the ends of that spectrum. Earlier layer-wise diagnostics showed where safety behavior concentrates but did not say whether a layer strengthens or weakens refusal; SLDR scales each layer's attention and feed-forward weights up and down, measures the normalized signed change in refusal counts, and reads a positive score as refusal-enhancing and a negative score as refusal-suppressing. Across the 32 layers of Llama3.1-8B-Instruct, adjacent layers 12 and 13 both carry large sensitivity values but with opposite signs, so targeting only positive safety layers would miss a complementary failure mode.

What to watch

A next step is to check whether the diagnosis still reliably picks the two spectrum endpoints on models with far more parameters, since the current experiments stay in the 7B to 8B range. Engineering teams running fine-tuning services can treat it as a post-fine-tuning repair stage: run the offline probe once to select layers, train the recovery adapter on a small alignment set, and decide per query whether to activate it. For those working on routing, the behavior of the malicious-score threshold on out-of-distribution inputs is worth testing further.

Whether the signed sensitivity spectrum keeps the same two-end structure on larger models and under more complex attacks still needs more settings to confirm. Routing relies on representation similarity, so an attacker who crafts inputs whose malicious score falls below the threshold could bypass conditional activation, a robustness question the text raises without resolving. Under low-resource-language encoding jailbreaks, harmful scores rise for every method, and SLDR is lowest but still residual, making the router's hardening against such distribution shifts a direction to watch.

Sources