The Multimodal Anonymizer: a fully local multi-agent AI system for medical data deidentification
Synopsis
The study developed and evaluated the Multimodal Anonymizer, a modular, locally deployable multi-agent framework integrating multimodal large language models, task-specific neural networks, and rule-based transformations; on benchmarks spanning text, tables, PDFs, imaging, metadata, filenames, audio, handwriting, and 3D imaging, its best local configuration (orchestrator Qwen3-VL-235B-A22B-Thinking) achieved 98.80% per-patient deidentification sensitivity (95%-CI 97.20; 100) and 99.60% critical clinical preservation (95%-CI 98.80; 100), reached 100% sensitivity and critical preservation on 250 local Charité partograms, and outperformed established tools across most modalities.
Interpretation
It introduces and validates a unified multi-agent deidentification framework that integrates MLLM contextual reasoning, task-specific neural networks (e.g., fine-tuned RetinaNet for defacing volumetric head scans, fine-tuned ResNet-UNet for face detection, fine-tuned TrOCR for handwriting) and rule-based transformations (regex date detection with a patient-specific ±3-year date shift) into a single pipeline with an iterative verification agent. Prior approaches were largely modality-specific and lacked contextual interpretation of identifiers embedded across data types; this work integrates three computational strategies into an orchestratable system covering text, tables, PDFs, imaging, metadata, filenames, audio, handwriting, and 3D imaging. Evaluated on 250 MIMIC-IV patients with injected synthetic PII and exact annotations plus modality-specific datasets and 250 Charité partograms; the best configuration reached 98.80% per-patient sensitivity (95%-CI 97.20; 100) and 99.82% per-PII sensitivity (95%-CI 99.76; 99.88), with all modalities at or above 98.30% sensitivity (lower 95%-CI bound); primary outcomes were manually reviewed by two independent experts.
Beyond deidentification sensitivity, it jointly quantifies clinical content preservation, reporting 99.60% critical clinical preservation (95%-CI 98.80; 100, per-patient) and 99.61% clinical preservation (95%-CI 99.51; 99.71, per-file), with a single instance of clinically critical information loss across all datasets (one active pharmaceutical ingredient redacted from a discharge note). Prior evaluations focused largely on sensitivity; this work treats disclosure risk and data utility as joint outcomes, aggregates at the patient level to reflect cumulative reidentification risk, and preserves longitudinal temporal structure through date shifting. Critical clinical preservation and clinical preservation were assessed by two independent domain experts with disagreements resolved by discussion; on local Charité data, critical clinical preservation was 100% per-patient and clinical preservation 99.97% (95%-CI 99.91; 100) per-file.
The system is deployed fully locally and plug-and-play, with code publicly available on GitHub and a browser-based interface where users configure a local MLLM endpoint, upload files, and download compressed outputs preserving original folder structures without programming expertise. It turns multimodal deidentification from a fragmented collection of modality-specific tools into a unified institutional workflow while meeting data residency and institutional control requirements. Open-source models were deployed locally on a cluster with 16 NVIDIA H200 GPUs, with GPT-5.2 accessed through a secured Azure deployment only as an upper-bound comparator; documentation, prompts, model settings, and package versions are provided.
The system outperformed existing tools on most modalities, and leading local models were similar to the proprietary GPT-5.2 in deidentification sensitivity while showing higher specificity. It achieved significantly higher deidentification sensitivity than Presidio, ORB-HD, and scitlab on tables, imaging, handwriting, German text, and GRASCCO, and supported metadata and filename deidentification and date shifting that comparators did not. Paired bootstrap comparisons showed higher table sensitivity than Presidio (Δ 4.61% ± 3.91; 5.30) and scitlab (Δ 27.18% ± 26.60; 27.75), and higher imaging sensitivity than Presidio (Δ 61.24% ± 57.86; 64.58) and scitlab (Δ 23.62% ± 23.08; 24.04); GPT-5.2 had the highest overall sensitivity but the lowest specificity (77.37%), significantly below GLM-4.6V (Δ -15.50% ± -17.41; -13.38).
Perspective
The results are intended for research and clinical data-governance settings that need to safely reuse multimodal clinical data within institutional environments, for institutions with local GPU resources able to deploy open-source MLLMs and follow deidentification standards such as the 18 HIPAA Safe Harbor categories; the framework preserves original file formats and folder structures and can optionally retain a linkage file connecting deidentified and original data, supporting secondary use within governed environments.
The benchmark relied substantially on MIMIC-IV with synthetic identifier injection, which enabled precise ground truth but may not fully capture the variability of naturally occurring identifiers in routine clinical care; generalizability across institutions, languages, specialties, and document types warrants further validation. Residual reidentification risk cannot be fully eliminated, particularly through linkage of rare clinical features or external data sources. Best performance required substantial computational resources. Data utility was inferred from preservation metrics rather than tested directly in downstream modeling tasks. In addition, the loaded text is a fast parse in which some figures and supplementary materials are not fully rendered, which may affect verification of individual result details.
