AURA combines mask-reconstruct anonymization with adversarial screening to reach the lowest agentic re-identification counts among non-DP methods and recover 6.4 percentage points more contextual utility than a prior LLM anonymizer at comparable privacy
Related research and updatesSynopsis
The work introduces AURA (Anonymization with Utility-Retention Adaptation), an LLM-powered mask-reconstruct framework that decouples privacy localization from utility-preserving reconstruction and selects candidates through adversarial privacy and utility-retention checks; evaluated on real-user interview transcripts with re-identification attacks carried out by web-search agents, adaptive-scope AURA yields the lowest agentic re-identification counts under each of three attacker models among non-DP methods, and at matched scope and backbone it retains more contextual utility than the prior LLM anonymizer (+6.4 pp unit-grid recovery) at comparable privacy.
Figure 5: Pareto front for privacy success versus code-fact recovery. Code-fact preservation remains comparatively high for several systems, including the anonymizer, showing that preserving thematic transcript content is easier than preserving the joint profile–code units used in the main-text utility-grid metric.
arXivInterpretation
It proposes AURA, an LLM-powered mask-reconstruct anonymization framework that separates privacy localization (which spans to mask) from utility-preserving reconstruction (generating replacement text) into two stages that can be optimized separately. Existing defenses either remove explicit identifiers, perturb text for formal privacy, or test rewritten text against non-web inference models; AURA explicitly targets the operating region between resistance to agentic web-search re-identification and utility retention, and bases candidate selection on adversarial privacy checks together with utility-retention checks. At the abstract level the framework design and selection mechanism are described, and evaluation on real-user interview transcripts is stated; implementation details, prompt design, and selection thresholds are not given in the abstract.
Under each of three attacker models, adaptive-scope AURA achieves the lowest agentic re-identification counts among non-DP methods. Prior LLM anonymizers were mainly tested against non-web inference models, leaving uncovered the threat model in which web-search agents cross-reference weak contextual cues into re-identification evidence. Evaluation uses real-user interview transcripts, attacks are carried out by web-search agents, and three attacker models are covered; the abstract reports a relative ordering (lowest counts) rather than specific count values.
At matched scope and backbone, AURA's mask-reconstruct design retains more contextual utility than the prior LLM anonymizer at comparable privacy, with a 6.4 percentage point higher joint contextual utility grid recovery. The comparison places utility and privacy under the same matched conditions, indicating a measurable utility-retention gain for the decoupled mask-reconstruct route over the prior LLM anonymization route. Utility evaluation rests on interviewee-profile facts, codebook facts, and the joint contextual utility grid; the abstract reports the +6.4 pp grid-recovery difference and states that privacy is comparable.
Perspective
The work targets researchers and data-governance practitioners who must control re-identification risk while preserving analytic value when releasing or sharing text; the setting is text rich in contextual detail (here, real-user interview transcripts), and the threat model is agentic re-identification with web search. Its value proposition is to offer selectable candidate rewrites within the operating region between resistance to agentic web-search re-identification and utility retention, using adversarial privacy checks and utility-retention checks as the selection basis; the adaptive-scope configuration is the key setting for achieving the lowest re-identification counts. On the utility side, applicability is measured through interviewee-profile facts, codebook facts, and the joint contextual utility grid.
The abstract does not give the sample size of interview transcripts, the specific settings of the three attacker models, the absolute re-identification counts, or the confidence intervals or variance behind the +6.4 pp figure, so the stability of the effect size across data scales is hard to judge. Utility is measured through interviewee-profile facts, codebook facts, and the joint contextual utility grid; semantic fidelity and downstream task performance beyond these measures remain open questions. How adaptive scope adjusts with text sensitivity, and how the method behaves on non-interview text or across languages, are not addressed in the abstract. In addition, this is a fast parse that does not include the figures and experimental details of the full text, so reproducing or comparing specific numbers should rely on the original.
