UDS measures LLM unlearning depth via activation patching, reporting top faithfulness and robustness across 150 models and 20 metrics
Synopsis
The work introduces the Unlearning Depth Score (UDS), which first uses a retain-model baseline to identify layers encoding target knowledge and then measures via activation patching how much of that knowledge the unlearned model has erased on a 0-1 scale; in a meta-evaluation across 20 metrics on 150 unlearned models spanning 8 methods, UDS is reported to achieve the highest faithfulness and robustness, and case studies show it uncovers residual knowledge obscured from observational metrics by representational shifts, with erasure depth varying across prompt types.
Figure 1: Overview of UDS for a single forget set example. (A) Stage 1 patches hidden states from M ret M_{\text{ret}} into M full M_{\text{full}} at each layer to measure how deeply the forget set knowledge is encoded. (B) Stage 2 repeats this with M unl M_{\text{unl}} as source to quantify how much encoded knowledge remains recoverable. (C) Stage 2 degradation is compared against Stage 1 at each layer to compute erasure ratios, which are weighted and aggregated into a single 0–1 score.
arXivInterpretation
It proposes UDS, a metric that quantifies the mechanistic depth of unlearning via activation patching: layers encoding the target knowledge are identified using a retain-model baseline, and the degree of erasure in the unlearned model is measured on a 0-1 scale. Existing output-level metrics fail to detect when knowledge remains recoverable from internal representations, while prior white-box studies reveal residual knowledge but often rely on auxiliary training or dataset-specific adaptations, leaving no generalizable metric; UDS offers a unified causal activation-patching measure. The abstract reports a meta-evaluation of 20 metrics on 150 unlearned models spanning 8 unlearning methods, with UDS achieving the highest faithfulness and robustness; specific values, datasets, and statistical details are not given in the loaded text.
Case studies show UDS uncovers residual knowledge that representational shifts hide from observational metrics, and that erasure depth varies across prompt types. This moves auditing from output behavior toward internal representations, indicating that output-only views may underestimate knowledge that has not truly been erased. Based on the case-study findings stated in the abstract; the text does not give the number of cases, the list of prompt types, or quantitative magnitudes.
The authors provide guidelines for integrating UDS into existing benchmarking frameworks and streamlining the evaluation pipeline, and release code and data. This gives the metric an operational path to adoption by existing evaluation pipelines rather than remaining a single-study artifact. The abstract explicitly mentions integration guidelines and a code/data link; the content of those guidelines is not expanded in the loaded text.
Perspective
The work targets settings where LLM unlearning must be audited, such as post-hoc evaluation for privacy protection and AI safety, and its users are research and engineering teams building or running unlearning benchmarks. UDS is set up so that a retain model serves as the baseline for locating layers that encode the target knowledge, after which activation patching measures the erased proportion in the unlearned model and yields a 0-1 depth score; it is designed to be embedded in existing benchmarking frameworks and to streamline the evaluation pipeline. The case studies indicate that applicability also depends on prompt type, since erasure depth varies across prompt types.
The loaded text is abstract-level, without specific metric values, dataset composition, prompt-type lists, or statistical-test details, so the margin by which UDS outperforms other metrics cannot be judged from the available material. The exact selection of the 20 metrics and the comparison conditions in the meta-evaluation, as well as how UDS behaves across model scales or domains, remain open questions a reader would need the full paper to resolve. In addition, the text does not explain the mechanism behind erasure depth varying across prompt types.
