Uncoded Clinical Features from Multilingual Electronic Health Records in Catalonia: Development and Validation Study
Synopsis
This study deployed the 3.8B-parameter open-weight small language model Phi4-mini (via Ollama) within an institutional firewall, combined with deterministic regular expression post-processing, to extract six uncoded urinary tract infection clinical features (fever, nitrites, leukocytes, lumbar pain, abdominal pain, and haematuria) from 15,498 Catalan/Spanish bilingual MEAP primary care narratives in the SIDIAP database in Catalonia, achieving 93.8% accuracy, 96.6% specificity, and 83.6% sensitivity in a double-blind clinician gold standard validation of 60 real patient records, and 87.2% accuracy, 99.5% specificity, and 99.2% positive predictive value in adversarial synthetic stress-testing of 720 notes, with no patient data leaving institutional servers.
Figure 1. Extraction pipeline scheme
medRxiv · Page 20Interpretation
Development and validation of a fully local, privacy-preserving hybrid extraction pipeline combining probabilistic small language model inference with deterministic regular expression rules for extracting uncoded clinical features from multilingual primary care narratives. Prior clinical NLP approaches predominantly relied on cloud-hosted commercial LLM APIs or BERT-based models fine-tuned for English-language EHRs; this study instead employs a 3.8B-parameter open-weight model (Phi4-mini) running locally within the institutional firewall, with prompt engineering and regex post-processing tailored to Catalan/Spanish bilingual and MEAP semi-structured formats. In double-blind clinician gold standard validation of 60 real records, overall accuracy was 93.8% (95% CI 90.6%-96.1%), specificity 96.6% (95% CI 93.6%-98.4%), and sensitivity 83.6% (95% CI 73.0%-91.2%); in adversarial synthetic stress-testing of 720 notes, accuracy was 87.2% (95% CI 84.6%-89.6%), specificity 99.5% (95% CI 98.1%-99.9%), and PPV 99.2% (95% CI 97.2%-99.9%).
Across 2,962 matched case-control patients, the pipeline identified 4,663 clinical feature occurrences, with cases progressing to acute pyelonephritis showing a higher feature burden than controls. The study applied extraction results to enrich predictive risk models for acute pyelonephritis within the ITUCAT observational study, demonstrating the usability of uncoded narrative data in real-world epidemiological research. Cases had a higher proportion of features than controls (71.3% vs 55.8%; standardized mean difference SMD = 0.326), particularly for fever (33.0% vs 9.0%; SMD = 0.616) and lumbar pain (29.0% vs 9.8%; SMD = 0.502).
A dual-validation strategy (real-world clinical gold standard and adversarial synthetic stress-testing) systematically assessed pipeline performance under edge cases including linguistic noise, negation traps, and rare synonyms. Most prior studies employed only a single validation approach; this study provides both clinical fidelity assessment and rigorous testing of edge-case failure modes, with an additional comparative analysis using Gemma3:4b on the same 60 records. Inter-reviewer agreement in real-world validation was 84.4% (Cohen's kappa), and overall Cohen's kappa between the two SLMs was 66.9 (47.4-86.4); Gemma3:4b showed overall accuracy of 84.9% (80.6-88.5), sensitivity of 97.3% (90.5-99.7), and specificity of 81.4% (76.2-85.9).
Perspective
The pipeline is intended for application within institutional firewalls, primarily for extracting features from Catalan and Spanish primary care MEAP narratives, with current validation restricted to the specific clinical context of urinary tract infections. Its modular design allows prompts and regular expressions to be adapted to new clinical domains or languages; the authors propose future generalization to areas such as cardiovascular and respiratory conditions, and potential use in routine management tasks including clinical documentation quality monitoring, patient profiling, or resource planning. For other healthcare systems facing similar privacy constraints (e.g., hospital frameworks), the pipeline offers a transferable template.
Readers may consider the following open questions: the reasoning capacity boundaries of a 3.8B-parameter model when handling complex negations, non-standard clinical shorthand, and narrative noise; the potentially conservative feature capture tendency arising from prioritizing regex outputs to maximize specificity; the selection bias potentially introduced by excluding 23 narratives (6.4% of the validation sample) due to persistent inter-annotator disagreement in real-world validation; the fact that the six extracted variables represent a mixture of symptoms, clinical signs, and urinalysis findings and should not be interpreted as equivalent clinical manifestations or independent diagnostic criteria for pyelonephritis; the high heterogeneity of MEAP documentation completeness depending on individual clinicians, time available per consultation, and overall visit volume, which may affect representativeness of extraction results; and the generalizability of the pipeline to clinical domains beyond urinary tract infections, which remains to be evaluated.
