Nullify achieves training-free LLM unlearning via null-space activation steering, raising TOFU forget10 FQ to 0.58 while holding MU at 0.62, matching Retrain
Synopsis
Nullify introduces a training-free, weight-preserving inference-time activation steering method for LLM unlearning: it builds a forget direction from IDK-contrastive activations and confines every correction to the subspace orthogonal to retained-knowledge activations via a null-space projector, matching or surpassing baselines in forget quality on TOFU and MUSE while keeping model utility near-lossless at a small fraction of their compute.
(a) Forget quality vs. model utility.
arXivInterpretation
Moving unlearning from parameter space to activation space and constraining it with a null-space projector decouples forgetting from utility preservation: the correction vanishes on retain queries by construction, so forgetting no longer has to be traded against capability. Prior gradient-based (GA, GradDiff), preference-optimization (NPO, SimNPO) and representation-level (RMU, LUNAR) methods either update the base weights or train auxiliary parameters, with retention enforced only by a soft loss term; Nullify is training-free, weight-preserving, and backed by a principled utility-preservation construction. Ablations isolate the two components: replacing the null-space projector with the identity leaves MU at only 0.360 even at the most conservative strength and collapses it as strength grows; replacing the IDK-contrastive direction with a random vector halves FQ to 0.466 while MU stays near 0.600, showing the projection preserves utility and the contrastive direction carries the forget signal.
The whole pipeline is closed-form and annotation-free: the per-layer operator is solved in closed form from one forward pass, the forget direction comes from the residual between memorized-continuation and abstention-template hidden states on the forget set itself, and steering strength is a runtime scalar. Zero trainable parameters and no backpropagation; offline work is one Gram SVD and one pseudoinverse, and at inference the per-prompt overhead is a single matrix-vector product per steered layer, with steered hidden states populating the KV cache so no per-token hook is needed. Complexity analysis gives the per-layer offline cost and an asymptotically negligible inference overhead; hyperparameter sweeps show FQ peaks inside a wide plateau and MU stays within 0.607-0.628 across all steering strengths, indicating utility preservation comes from the construction rather than tuning.
On TOFU and MUSE-News it delivers the strongest forget-utility balance: TOFU forget10 FQ rises from SimNPO's 0.45 to 0.58 with MU held at 0.62, equal to Retrain; on MUSE-News PrivLeak is 5.18, an order of magnitude closer to the ideal 0 than any baseline. Baselines fail in opposite directions: NPO (108.91), GradDiff (108.12) and SimNPO (72.93) leak heavily, Task Vector (100.00) over-corrects past the retrained reference, LUNAR (99.80) barely moves from the un-unlearned level, and GA collapses with KM-D at zero. On TOFU, Nullify's World Facts ROUGE-L of 0.90 and Retain-set 0.95 track the Original model (0.91/0.98); on MUSE-News forget-set KM-D drops from the Original's 0.63 to 0.27 (Retrain 0.33) while retain-set KM-D stays at 0.51 (Retrain 0.54).
Deployment cost drops sharply: Nullify runs in pure inference memory at 14 GB and finishes one configuration in 5.2 minutes, whereas full-finetune baselines carry a 6.6 B-parameter Adam state, saturate 79 GB, and take 13-58 minutes. In the forget-quality versus wall-clock plane, GA, GradDiff, IDK, NPO and LUNAR are each both slower and weaker; only SimNPO and the Retrain oracle reach higher forget quality, paying 5.3x and 57.7x Nullify's wall-clock respectively. End-to-end timing on identical hardware (a single NVIDIA H800 80 GB) under the TOFU forget05 / Llama-2-7B-chat-hf setting, including training and default TOFU evaluation; because steering strength does not enter the closed-form preparation, adapting to a new forget request needs only one inference-time evaluation per strength value.
Perspective
The work targets the serving-time inference-intervention setting: while the operator is active it suppresses access to memorized answers rather than permanently altering the underlying weights, so it suits service environments that need to mask specific memorized content on request. It lets privacy-, copyright- and safety-driven deletion requests be fulfilled without touching the base model, which is especially relevant to providers who must answer deletion requests frequently but cannot afford retraining. Cross-family replication shows the construction and layer-selection procedure transfer to Qwen3-8B, though the observed forgetting-utility tradeoff differs across the two backbones tested. The authors name multimodal architectures and long-context unlearning as future directions.
As an inference-time intervention, suppression depends on the operator being active and the weights themselves are not permanently altered, so readers should watch white-box settings in which the operator can be inspected, modified, or disabled, as well as adaptive attacks, both of which the authors place outside the current scope. The robustness evaluation covers fixed-template black-box extraction and nearby free-form reformulations; leakage under paraphrasing (0.218) is below the Original (0.478) but above Retrain (0.128), so how far it transfers to more distant query formats remains open. Intervention layers and steering strength are currently selected per benchmark, and automatic configuration across architectures and deployment settings is unresolved. In addition, some numeric values in the loaded text appear as placeholders, so exact table figures should be checked against the original when precise reproduction is needed.
