Skip to main content
Back to timeline
Journal of Medical Internet ResearchSource publication:

Anonymization of Portuguese Clinical Notes Using Large Language Models and Quantum-Enhanced Hybrid Architectures: Comparative Evaluation Study

Synopsis

This study built a gold-standard corpus of 1000 Portuguese outpatient clinical notes manually annotated by 5 trained researchers for 5 protected-entity categories (patient names, dates, identifiers, organizations, and geographic locations) and, on a held-out test set of 500 notes, compared two stand-alone LLMs (Llama-3.1-8B-instruct and Llama-3.3-70B-instruct) with two quantum-enhanced hybrid models (Dynex-QML with 8B and 70B base models, using QUBO formulations to transform the final attention layer into a global constraint satisfaction problem solved by neuromorphic quantum annealing); the quantum-enhanced Dynex-QML-70B achieved the highest macro-F1 of 0.855 (95% CI 0.823-0.880), above stand-alone Llama-3.3-70B (0.726), Dynex-QML-8B (0.733), and Llama-3.1-8B (0.

Source-provided article image: Anonymization of Portuguese Clinical Notes Using Large Language Models and Quantum-Enhanced Hybrid Architectures: Comparative Evaluation Study.
Figure 1. ·

Comparison of entity-level F 1 -scores across 4 anonymization models, Llama 3.1 8B, Llama 3.3 70B, Dynex-QML 8B, and Dynex-QML 70B, evaluated on 500 held-out Portuguese outpatient clinical notes from a tertiary-care academic hospital in Brazil. The evaluated protected entity categories included DATE, ID, LOCAL, NAME, and ORGANIZATION. Error bars indicate aggregate multinomial bootstrap 95% CIs. Brackets indicate pairwise comparisons. All starred comparisons had P <.001, except Llama 3.3 70B vs Dynex-QML 70B for LOCAL ( P =.03).

PubMed

Interpretation

The study provides an annotated Portuguese clinical-note corpus and a multi-entity anonymization benchmark covering five protected-entity categories: names, dates, identifiers, organizations, and geographic locations. Prior work has largely focused on English or on rule-based and machine learning methods; this work builds a gold standard from 1000 outpatient notes with 5 trained annotators and evaluates on 500 held-out notes using precision, recall, and F1. Corpus size and annotation procedure are explicitly stated, and evaluation uses a held-out test set with confidence intervals, making it a reproducible benchmark.

Quantum-enhanced hybrid architectures outperform same-size stand-alone LLMs on anonymization accuracy, with Dynex-QML-70B reaching a macro-F1 of 0.855, an improvement of 0.128 over stand-alone Llama-3.3-70B (95% CI 0.091-0.163; P < .001). This comparison is among the first to combine quantum optimization (QUBO formulation, neuromorphic quantum annealing) with an LLM attention layer for Portuguese clinical text and to directly contrast it with stand-alone models. The difference is tested with an empirical two-sided bootstrap, P < .001, and most entity-level within-size comparisons are also statistically significant, with the exception of the ORGANIZATION comparison between Dynex-QML (Llama 8B) and Llama 3.1 8B.

The quantum-enhanced model improves accuracy without sacrificing efficiency: Dynex-QML-70B processed each note in 7.97 seconds end-to-end versus 8.52 seconds for stand-alone Llama-3.3-70B, with a paired note-level bootstrap showing a mean reduction of 0.55 seconds (SD 2.36; 95% CI -0.75 to -0.34; P < .001). Prior evaluations of anonymization methods often emphasize accuracy alone; this work quantifies end-to-end processing time alongside accuracy, offering a joint performance-efficiency comparison. Timing is reported as per-note means with confidence intervals and compared via a paired note-level bootstrap, giving explicit statistical inference.

The gains of the quantum-enhanced architecture appear mainly as reductions in false-positive rates while preserving high sensitivity. The paper attributes the improvement to fewer false positives rather than lower recall, indicating that the constraint-satisfaction optimization suppresses over-labeling. This conclusion rests on the combined precision, recall, and F1 results, supported by macro-F1 and entity-level comparisons reported in the text.

Perspective

This work targets Portuguese outpatient clinical notes and five protected-entity categories (names, dates, identifiers, organizations, geographic locations), evaluated on a held-out test set; its findings support using quantum-enhanced hybrid architectures in deidentification settings that require high sensitivity while controlling false positives, and provide a starting point for validation in other languages, other entity types, and real deployment settings.

A careful reader would still watch for the stability of the quantum-enhanced component across different hardware and annealing conditions, performance in other languages and entity types, and details such as inter-annotator agreement; the loaded text is abstract-level and does not include figures or full methodological detail, so specific values and implementation details on these points would require consulting the original article.

Sources