Skip to main content
Back to timeline
arXivSource publication:

Needle removes LLM backdoors with one weight-orthogonalisation edit, cutting code-injection attack success to zero while keeping capability and safety loss lowest

Synopsis

The work proposes Needle, a training-free backdoor removal method: given an identified trigger, it estimates a backdoor direction and a refusal subspace from activation vectors, then applies sequential weight orthogonalisation, achieving the lowest mean attack success rate among the evaluated defences across Gemma-3-4B-IT, Qwen3-4B-Instruct-2507 and Gemma-3-12B-IT, fully removing a code-injection attack while producing the lowest KL divergence and minimal capability and safety change.

AI-generated editorial illustration: Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation

Interpretation

The authors find that backdoor behaviour can largely be captured by a single direction in activation space, and that this direction is highly correlated with directions mediating refusal, with much of the overlap lying outside the single primary refusal direction. Prior backdoor removal work was developed mainly for classification models or relies on clean reference models and inference-time intervention; here the backdoor is characterised as a locatable activation direction and its overlap with refusal representations is quantified. Based on experiments across two model families, three attack behaviours and two trigger types, with layer-wise cosine similarity showing low similarity to the primary refusal direction for sentiment steering and code injection that rises at rank four, while targeted refusal is already high at rank one.

Needle applies closed-form weight edits satisfying two constraints at once, removing the backdoor projection while preserving the projection onto the refusal subspace, with sequential layer-wise correction to compensate for drift caused by earlier edits. Unlike defences that replace tokens or steer activations at inference time, Needle permanently edits weights and requires neither a clean reference model nor the original poisoned training data. The method orthogonalises attention output and MLP down-projection matrices, fits the correction with ridge regression, and is compared in the appendix against minimum-norm least squares and LASSO with small differences (on the order of at most 1 point in ASR).

In evaluation Needle achieves the lowest mean attack success rate among the compared defences while also showing the lowest distribution shift and the smallest capability and safety cost. Fine-tuning-based defences tend to trade unreliable removal against capability and safety loss; Needle reports a combination that is favourable on removal and side effects together. Mean ASR falls on both Gemma and Qwen, code-injection ASR reaches 0 while the best baseline retains a clear residue, and average relative capability loss and KL divergence are lower than the fine-tuning baselines; targeted refusal is the hardest case, where Needle's ASR is highest.

Ablations show that preserving the refusal subspace and applying sequential correction both reduce the safety cost, and that editing only middle-to-late layers is preferable to editing early layers. The natural baseline of directly ablating the backdoor direction is decomposed into separable design choices, with the trade-offs of each choice reported across ASR, capability and harmful-response rate. On the BadNet sentiment steering attack, directly ablating the backdoor direction sharply raises harmful responses; preserving a single refusal direction, then extending to the refusal subspace, then adding sequential correction progressively improves safety, while editing early layers removes the backdoor completely but increases capability loss and KL.

Perspective

The result targets defence settings where the trigger and insertion rule have already been identified, and applies to open-weight models where the defender has white-box access but not the original poisoned data or a clean reference model; the authors position Needle as a targeted removal step complementary to trigger detection and reconstruction work, and release an open-source implementation. For practitioners this means that, given a reliable detection step, a single weight edit can replace inference-time intervention or broad fine-tuning, preserving existing capability and refusal behaviour in computationally constrained settings.

The authors state as a limitation that the experiments cover a set of synthetic data-poisoning attacks with explicit trigger-behaviour associations, and do not capture naturally occurring spurious behaviours or attacks deliberately designed to evade representation-based removal; whether removal persists under subsequent fine-tuning or re-poisoning also remains to be tested. In addition, Needle's ASR is highest on the targeted refusal attack, which the authors explain by the backdoor target itself being a refusal, creating a trade-off with the refusal behaviour that must be preserved. This evidence bundle is the full text, but some tables and appendix values appear as placeholders in the text, so exact per-item numbers should be checked against the original.

Sources