Skip to main content
Back to timeline
medRxivSource publication:

Task-dependent model selection for structured extraction from multilingual non-English clinical records

Synopsis

Across 193,101 Russian- and Kazakh-language stroke discharge summaries, with test cohorts of 332 section cases, 149 medication cases (2,475 reference records), and 191 laboratory cases, this study compared a multilingual encoder, locally fine-tuned Qwen3-4B models, and zero-shot GPT-5.5 on entity detection versus complete-record assembly, finding section F1 of 0.919/0.926/0.932, GPT-5.5 leading drug-name detection (0.966 versus 0.940) while Qwen led normalized medication recovery (0.381 versus 0.311; difference 0.070, 95% CI 0.026–0.114) and laboratory quintuple F1 (0.892 versus 0.822), and showing that moving from curated sections to a raw-document cascade reduced medication recovery from 0.377 to 0.204 (single window) and 0.246 (all blocks) and laboratory quintuple F1 from 0.898 to 0.

AI-generated editorial illustration: Task-dependent model selection for structured extraction from multilingual non-English clinical records

Interpretation

On the detection-oriented section task, the compact multilingual encoder was close to the generative models: section-span F1 was 0.919 for the encoder, 0.926 for GPT-5.5, and 0.932 for Qwen, with a local-minus-cloud difference of 0.006 (95% CI −0.010 to 0.021; Holm p=0.804). Prior clinical extraction comparisons have concentrated on English and recurring institutional sources and have rarely separated detection from assembly within one corpus; this study provides same-task, same-scoring contrasts of three model families on one Russian–Kazakh stroke corpus. Based on a 332-case consensus test set and 1,019 reference spans, with document-bootstrap intervals and Holm adjustment; the interval spans zero, so the difference on this task is not clear.

Model ranking reversed with the output requirement: GPT-5.5 led drug-name detection (0.966 versus 0.940; local-minus-cloud −0.026, Holm p=0.011), whereas locally fine-tuned Qwen led normalized seven-field complete-record recovery (0.381 versus 0.311; difference +0.070, 95% CI 0.026–0.114, Holm p=0.002) and laboratory quintuple F1 (0.892 versus 0.822; difference +0.070, 95% CI 0.047–0.094). The study explicitly separates identifying mentions from linking their attributes into complete records and shows that the same models can rank differently on the two objectives, which prior clinical NLP comparisons centered on named-entity recognition have less often shown. Medication results rest on 149 cases and 2,475 reference records, laboratory results on 191 cases; raw-field recovery on the same cohort was 0.080/0.173/0.284 (encoder/cloud/local), showing sensitivity to normalization rules.

Replacing curated sections with a raw-document cascade reduced record-level performance: medication recovery fell from 0.377 in the matched curated control to 0.204 for the single-window cascade (change −0.173, 95% CI −0.216 to −0.131) and 0.246 for all blocks (change −0.131, 95% CI −0.180 to −0.084), and laboratory quintuple F1 fell from 0.898 to 0.811 (change −0.087, 95% CI −0.111 to −0.065). By placing component benchmarks alongside end-to-end raw-document evaluation, the study quantifies how upstream section extraction and joining affect downstream record assembly, rather than reporting only extraction from prepared inputs. Based on matched controls and cascades over 149 medication and 191 laboratory cases with document-bootstrap intervals; error inspection also found drug-name omissions when the name remained in the supplied text, indicating both upstream content loss and downstream generation errors.

Compared with a second annotator, the generative models agreed with reference annotations on normalized complete medication records more closely than the original annotator did: A1 versus A2 was 0.178, with model-minus-human differences of +0.175 for cloud (Holm p<.001) and +0.172 for local (Holm p<.001); on 50 laboratory reports, human quintuple agreement was 0.902 and the local model reached 0.898, a paired difference of −0.003 (95% CI −0.020 to +0.013). The study uses annotator agreement as an empirical reference for interpreting model scores rather than reporting absolute metrics alone, bringing annotation conventions and normalization into the discussion of measured performance. Fifty held-out paired cases per task, using a common second-annotator reference and paired document-bootstrap intervals; the sample is limited, which the authors also list as part of the evaluation scope.

Perspective

The results apply to retrospective, offline secondary use of Russian- and Kazakh-language stroke discharge summaries from one health system, and to institutions weighing detection-oriented tasks against complete-record assembly; they support choosing an approach according to the required record detail and local processing arrangements, and they indicate that the full pathway from raw documents to records should be evaluated rather than only extraction from prepared inputs.

Readers may still watch that the section test cohort predominantly contained documents with all three target labels, giving limited evidence on documents with missing sections; human comparisons used 50 reports per task; incomplete patient-level linkage leaves possible overlap between records from the same patient; DAPT included unlabeled test text, and the non-random exposed and unexposed groups contained 179 and 153 cases with encoder F1 rounding to 0.919 in both, so an exposure effect cannot be separated from group differences; the separate contributions of model family, capacity, training exposure, prompting, annotation ambiguity, and normalization remain unresolved; laboratory macro-F1 is complicated by category aliases and empty-support categories; runtime and cloud charges cover different hardware and processing scopes and exclude governance and human-review costs; and clinical workflow effectiveness, safety, and patient benefit require separate evaluation.

Sources