Clinical Code Mapping with LLM Tool Use: A Pilot for Automated Data Extraction of Medication and Diagnosis Information from Unstructured Clinical Notes
Synopsis
This pilot study anonymized 35 German doctor's notes from five patients, built one pipeline for medication extraction and mapping and two for diagnoses (one RAG-based and one agentic AI), ran them with three open-weight LLMs on a local GPU-PC, and found that medication name extraction reached an F1 of 0.95 and medication mapping 0.78, while diagnosis coding did not exceed an F1 of 0.12 and broad-category mapping reached 0.18, leading the authors to conclude that LLMs are suitable for medication information extraction for research databases but that current state-of-the-art open-weight models are not accurate enough for a clinical setting where patient treatment would depend on LLM performance.
Interpretation
The study built and compared LLM-based structured extraction pipelines for German clinical notes: one for medication extraction and mapping, and two for diagnoses, with the diagnosis part comparing a RAG-based approach against an agentic AI approach. Conventional structured data extraction often relies on manual curation, which is prohibitively labor-intensive; this work brings open-weight LLMs and tool-use pipelines to non-English (German) clinical notes and directly compares RAG and agentic routes for diagnoses. Based on 35 anonymized German doctor's notes from five patients, run with three open-weight LLMs on a local GPU-PC, in a small pilot design.
Medication-related tasks performed clearly better than diagnoses: medication mapping reached an F1 of up to 0.78, and up to 0.95 when considering trivial name extraction only. Provides a quantified comparison of task difficulty on the same set of notes, showing the performance gap between extraction and code mapping. Reports specific F1 values, and the authors note that for trivial name extraction of medications every encountered mistake is explainable.
Diagnosis coding performed poorly: F1 did not exceed 0.12, and mapping to the broad category raised F1 to 0.18. Quantifies the actual level of the harder diagnosis coding task under current open-weight models and shows that relaxing mapping granularity yields only limited improvement. Reports specific F1 values on 35 notes from five patients, a pilot-scale sample.
The authors conclude that LLMs excel at extraction and are suitable for medication information extraction from clinical notes for use in research databases, but that in a clinical setting where patient treatment would depend on LLM performance, current state-of-the-art open-weight models are not accurate enough. Translates the performance differences into an explicit boundary between research use and clinical decision use. The conclusion rests on the reported F1 results and the authors' discussion of task-nature limitations, noting that due to the nature of the task it is infeasible to expect a perfect score of 1 in any coding scenario.
Perspective
The results are intended for research settings that use open-weight LLMs to extract structured information from German clinical notes, especially medication name extraction and medication mapping for building research databases; the authors explicitly state that in a clinical setting where patient treatment would depend on LLM performance, current state-of-the-art open-weight models are not accurate enough. For diagnosis coding, even broad-category mapping reached only an F1 of 0.18, so that route currently fits better as an exploratory direction.
This is a pilot-scale effort involving only 35 German notes from five patients and three open-weight models, so how the high medication F1 and low diagnosis F1 hold up in larger samples, other languages, or other hospital settings remains to be observed; the authors also mention further problems in LLM output and parsing, whose scope of impact awaits follow-up work.
