Public articles linked to the same research event.
medRxiv This study deployed the 3.8B-parameter open-weight small language model Phi4-mini (via Ollama) within an institutional firewall, combined with deterministic regular expression post-processing, to extract six uncoded urinary tract infection clinical features (fever, nitrites, leukocytes, lumbar pain, abdominal pain, and haematuria) from 15,498 Catalan/Spanish bilingual MEAP primary care narratives in the SIDIAP database in Catalonia, achieving 93.8% accuracy, 96.6% specificity, and 83.6% sensitivity in a double-blind clinician gold standard validation of 60 real patient records, and 87.2% accuracy, 99.5% specificity, and 99.2% positive predictive value in adversarial synthetic stress-testing of 720 notes, with no patient data leaving institutional servers.
This study deployed the 3.8B-parameter open-weight small language model Phi4-mini (via Ollama) within an institutional firewall, combined with deterministic regular expression post-processing, to extract six uncoded urinary tract infection clinical features (fever, nitrites, leukocytes, lumbar pain, abdominal pain, and haematuria) from 15,498 Catalan/Spanish bilingual MEAP primary care narratives in the SIDIAP database in Catalonia, achieving 93.8% accuracy, 96.6% specificity, and 83.6% sensitivity in a double-blind clinician gold standard validation of 60 real patient records, and 87.2% accuracy, 99.5% specificity, and 99.2% positive predictive value in adversarial synthetic stress-testing of 720 notes, with no patient data leaving institutional servers.
This study deployed the 3.8B-parameter open-weight small language model Phi4-mini (via Ollama) within an institutional firewall, combined with deterministic regular expression post-processing, to extract six uncoded urinary tract infection clinical features (fever, nitrites, leukocytes, lumbar pain, abdominal pain, and haematuria) from 15,498 Catalan/Spanish bilingual MEAP primary care narratives in the SIDIAP database in Catalonia, achieving 93.8% accuracy, 96.6% specificity, and 83.6% sensitivity in a double-blind clinician gold standard validation of 60 real patient records, and 87.2% accuracy, 99.5% specificity, and 99.2% positive predictive value in adversarial synthetic stress-testing of 720 notes, with no patient data leaving institutional servers.
This study deployed the 3.8B-parameter open-weight small language model Phi4-mini (via Ollama) within an institutional firewall, combined with deterministic regular expression post-processing, to extract six uncoded urinary tract infection clinical features (fever, nitrites, leukocytes, lumbar pain, abdominal pain, and haematuria) from 15,498 Catalan/Spanish bilingual MEAP primary care narratives in the SIDIAP database in Catalonia, achieving 93.8% accuracy, 96.6% specificity, and 83.6% sensitivity in a double-blind clinician gold standard validation of 60 real patient records, and 87.2% accuracy, 99.5% specificity, and 99.2% positive predictive value in adversarial synthetic stress-testing of 720 notes, with no patient data leaving institutional servers.