Representation geometry explains why forgotten knowledge is easily relearned, motivating an MCU unlearning method that targets minor components
Synopsis
This work studies, from a representation geometry perspective, why unlearned models rapidly recover forgotten knowledge under relearning attacks, finding that existing unlearning methods mainly optimize along dominant components while leaving minor components largely unchanged, that modifications in dominant components are easily reversed during relearning whereas minor components resist such reversal, and proposes Minor Component Unlearning (MCU), which explicitly targets minor components and is validated on three datasets to substantially improve resistance to relearning attacks.
Figure 1: Left: Retraining-on- T T (RTT) attack evaluation: the forget set is split into T T and V V ; after unlearning on T ∪ V T\cup V , the attacker fine-tunes on T T and measures recovery on V V . Middle: Naive methods and SAM separate forget/retain representations mainly along dominant components ( DC ), which relearning easily reverses; MCU additionally separates them along minor components ( MC ), whose changes are largely preserved post-attack. Right: On WMDP-Cyber, MCU yields markedly lower post-attack accuracy while maintaining utility.
arXivInterpretation
The work identifies the mechanism by which unlearned models rapidly recover forgotten knowledge under relearning attacks: existing unlearning methods predominantly optimize along dominant components of representations, leaving minor components largely unchanged. Relative to prior work that reports this fragility as a phenomenon, this work offers a mechanistic explanation from representation geometry, attributing the fragility to optimization concentrated in dominant components. Based on analysis of the representation geometry of existing unlearning methods, the authors report the observation that dominant components are optimized while minor components remain largely unchanged.
The work finds that modifications in dominant components are easily reversed during relearning attacks, enabling rapid knowledge recovery, whereas minor components exhibit stronger resistance to such reversal. This contrast turns which directions are more robust into an actionable design principle rather than a mere description of unlearning failure. Obtained from observations of representation changes under relearning attacks, supported by a theoretical analysis from the spectral structure of representations that explains both observations.
The work proposes Minor Component Unlearning (MCU), which explicitly targets minor components in representations, concentrating unlearning effects in inherently more robust directions. Unlike existing methods that optimize along dominant components, MCU shifts the unlearning target to minor components to improve resistance to relearning attacks. The method design builds directly on the preceding mechanistic observations and the spectral-structure theoretical analysis.
Experiments on three datasets validate that MCU substantially improves resistance to relearning attacks. It translates mechanistic insight into a measurable robustness gain, offering a new methodological path for unlearning in open-weight model settings. Validated through experiments on three datasets; the abstract reports substantially improved resistance to relearning attacks.
Perspective
The work targets settings where specific data influences must be removed from a pretrained model and model weights may be public, such as privacy-, copyright-, and safety-related unlearning needs. Its intended readers are researchers and engineers working on LLM unlearning, model safety, and representation analysis. The method applies to robustness evaluation under a relearning-attack threat model and has been validated on three datasets; the abstract does not specify the datasets, model scales, or deployment conditions, so practical scope should be judged from the full text.
The abstract does not provide specific evaluation metrics, attack strengths, baseline comparison numbers, or the names of the three datasets, nor does it state the assumptions and applicability conditions of the theoretical analysis; these details affect judgment of the magnitude of the robustness gain. In addition, whether the observation that minor components resist reversal holds across model scales and architectures, and how MCU affects general model capabilities, remain to be confirmed with the full-text experiments and follow-up work.
