Skip to main content
Back to timeline
arXivSource publication:

TIDE 2.0 separates an exchangeable recognizer from a keyed anonymizer, reaching span-level recall 0.88 and 0.77 on two institutions' corpora

Synopsis

TIDE 2.0 is an MIT-licensed, institution-hosted engine for clinical-note de-identification that separates an interchangeable PHI recognizer from a keyed, table-free anonymizer: the anonymizer applies a per-patient interval-preserving date shift, format-preserving encryption and deterministic keyed selection so that the same value receives the same surrogate across a patient's notes under one key and releases under a new key cannot be linked to earlier ones; the default configuration, a distilled recognizer plus a regex layer, reached span-level recall 0.88 at precision 0.88 on Stanford's SHIELD corpus and 0.77 at 0.87 on a second institution's i2b2 2014 corpus.

Source-provided article image: TIDE 2.0: an open, model-agnostic engine for keyed de-identification of clinical notes
Figure 1 ·

Figure 1: One sentence through the pipeline. The recognizer stage detects and types PHI spans and caches them before any text is altered. A reviewer can therefore inspect the detected spans before the anonymizer runs. The anonymizer stage applies a keyed strategy to each category. Both dates move by the same per-patient offset, so the fourteen-day interval between them is preserved. Fig. 4 shows the component-level architecture.

arXiv

Interpretation

The engine splits recognition and anonymization into two independently replaceable stages that communicate through a cached span representation, so a site can adopt a recognizer now and swap it later without re-anonymizing notes already processed; caching spans before any text is altered makes recognition auditable. Earlier open clinical de-identifiers coupled recognizer to anonymizer, or processed one note at a time with no execution layer; TIDE 2.0 opens the recognizer slot to any Hugging Face token classifier, an LLM recognizer, regular expressions and an Aho-Corasick gazetteer behind a Presidio-compatible interface. The architecture and interfaces are described in full in Methods; the paper benchmarks on a single commodity GPU and states that it does not report a warehouse-scale deployment.

The anonymizer is keyed and table-free: dates shift by a per-patient offset that preserves every interval, numeric and alphanumeric identifiers are encrypted with FF3 format-preserving encryption, and names and locations use deterministic keyed selection; the same value receives the same surrogate everywhere under one key, and a release under a new key cannot be linked to earlier releases. The authors state that few open tools generate surrogates that are both realistic and consistent across a patient's notes; TIDE 1.0 redacted most structured identifiers or replaced them with random values. On a seeded synthetic stress set of 2,000 distinct values per keyed category, determinism, cross-note consistency, remapping under a rotated key and interval preservation each held on 100.0% of trials; collisions are not guaranteed a priori: 3-5 digit identifiers shared a surrogate 34.8%-36.6% of the time, and name collisions were 0.4% among 2,000 names and 3.3% among 20,000.

The default configuration (TIDE2-Sentry plus a regex layer) reached micro-averaged recall 0.88 at precision 0.88 on SHIELD and 0.77 at 0.87 on i2b2 2014; the cross-institution recall drop of 0.11 was driven 56% by HOSPITAL, ID and PHONE together and 33% by DATE. The paper makes span-level recall the primary outcome rather than F1, on the rationale that a missed identifier is a reportable disclosure while a false positive costs a clinical word, and reports per-category results alongside the aggregates. Two gold-annotated corpora from two institutions (SHIELD: 1,381 notes, 10,229 spans; i2b2 2014 after crosswalk: 1,304 notes, 27,298 spans), with 95% bootstrap intervals over 2,000 document-level resamples; the authors state they did not analyse the missed spans, so the causes of the drops are hypotheses.

TIDE2-Sentry is a DeBERTa-v3-large token classifier distilled from Gemini 2.5 Flash over 13,000 unannotated notes; the student reached macro-averaged recall 0.81 against 0.90 for the teacher and had lower recall in every category. The paper quantifies distillation fidelity and shows that LOCATION is the weakest category for both teacher (0.66) and student (0.55), indicating the low recall there is shared with the teacher rather than introduced by distillation alone. The teacher was scored on all 1,381 SHIELD notes; after Bonferroni correction across the 9 categories, every teacher-student recall difference except WEB remained significant at 0.05, with the ID difference close to that threshold.

Perspective

The work targets institutions that need to turn clinical notes into research corpora on hardware they own, especially teams concerned with longitudinal analysis across notes and unlinkability between releases. The engine is designed to read from and write back to warehouse note tables, but the paper reports only a benchmark on one commodity GPU and does not report a warehouse-scale deployment. The recognizer slot is open, so any Hugging Face token classifier, LLM recognizer, or regex and Aho-Corasick matcher can fill it, letting a site adopt TIDE2-Sentry now and switch to a better model later without re-anonymizing notes already processed. Anonymizer keys are isolated per study or tenant, and a new key deterministically re-maps every surrogate so releases under different keys cannot be linked. The paper also provides a per-category decomposition of the cross-institution recall drop, which can indicate which categories need local gazetteers or format support when moving to a new site.

The paper did not run a downstream phenotyping or temporal-reasoning experiment on the de-identified corpus, so utility preservation is argued by construction and verified structurally, and downstream equivalence is not demonstrated. Throughput comes from single runs on one machine, the CPU-only path is an untuned single-actor configuration, and the study did not run on more than one node, measure overlap between stages, or test recovery from a failed worker. The student was trained on 13,000 teacher-generated silver labels, and the noise in these labels against gold was not characterised, so the effect of the teacher's systematic errors on the student is unquantified; the distillation notes are disjoint from both corpora at the note level, but patient-level disjointness from SHIELD was not verified. The AIMI recognizers were also trained on i2b2 2014, so that corpus is not held out for them and the paper does not use it to rank the two. Identifiers in i2b2 are surrogates, so recall there may differ from recall on real identifiers. Crosswalk decisions were made with access to both corpora, and without boundary alignment i2b2 micro-averaged recall is 0.69 rather than 0.77. The six external taggers are off-distribution on clinical notes and serve as a robustness probe. No LLM recognizer comparison is reported, and prompt development was carried out for the teacher model only. The anonymizer checks are unit-level tests on synthetic values; no linkage or re-identification attack was run on a release, and it was not tested whether surrogates make missed identifiers harder to find; membership inference is out of scope. Keyed, on-premises surrogate replacement lowers re-identification risk without eliminating it, offers no formal privacy guarantee of the kind defined by differential privacy, and the paper makes no claim that the output meets HIPAA Safe Harbor or k-anonymity.

Sources