Public articles linked to the same research event.
Studies in health technology and informatics This pilot study anonymized 35 German doctor's notes from five patients, built one pipeline for medication extraction and mapping and two for diagnoses (one RAG-based and one agentic AI), ran them with three open-weight LLMs on a local GPU-PC, and found that medication name extraction reached an F1 of 0.95 and medication mapping 0.78, while diagnosis coding did not exceed an F1 of 0.12 and broad-category mapping reached 0.18, leading the authors to conclude that LLMs are suitable for medication information extraction for research databases but that current state-of-the-art open-weight models are not accurate enough for a clinical setting where patient treatment would depend on LLM performance.
This pilot study anonymized 35 German doctor's notes from five patients, built one pipeline for medication extraction and mapping and two for diagnoses (one RAG-based and one agentic AI), ran them with three open-weight LLMs on a local GPU-PC, and found that medication name extraction reached an F1 of 0.95 and medication mapping 0.78, while diagnosis coding did not exceed an F1 of 0.12 and broad-category mapping reached 0.18, leading the authors to conclude that LLMs are suitable for medication information extraction for research databases but that current state-of-the-art open-weight models are not accurate enough for a clinical setting where patient treatment would depend on LLM performance.
This pilot study anonymized 35 German doctor's notes from five patients, built one pipeline for medication extraction and mapping and two for diagnoses (one RAG-based and one agentic AI), ran them with three open-weight LLMs on a local GPU-PC, and found that medication name extraction reached an F1 of 0.95 and medication mapping 0.78, while diagnosis coding did not exceed an F1 of 0.12 and broad-category mapping reached 0.18, leading the authors to conclude that LLMs are suitable for medication information extraction for research databases but that current state-of-the-art open-weight models are not accurate enough for a clinical setting where patient treatment would depend on LLM performance.
This pilot study anonymized 35 German doctor's notes from five patients, built one pipeline for medication extraction and mapping and two for diagnoses (one RAG-based and one agentic AI), ran them with three open-weight LLMs on a local GPU-PC, and found that medication name extraction reached an F1 of 0.95 and medication mapping 0.78, while diagnosis coding did not exceed an F1 of 0.12 and broad-category mapping reached 0.18, leading the authors to conclude that LLMs are suitable for medication information extraction for research databases but that current state-of-the-art open-weight models are not accurate enough for a clinical setting where patient treatment would depend on LLM performance.