Skip to main content
Back to timeline
medRxivSource publication:

Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes

Synopsis

Using 2,704 colorectal cancer discharge notes from MIMIC-IV and a 46-symptom inventory derived from the MSAS and EORTC QLQ-CR29, this study benchmarked dictionary-based rule matching, pretrained clinical NER, zero-shot Claude Haiku and Gemini 3.5 Flash, and two hybrid variants (LLM output plus post-hoc rule-based negation filtering) against a 200-note two-rater adjudicated gold standard, finding that Gemini 3.5 Flash performed best (Macro F1=0.70, Micro F1=0.86, Macro Precision=0.74), followed by Claude Haiku (Macro F1=0.63, Macro Recall=0.71), both substantially outperforming rule-based (Macro F1=0.44) and NER (Macro F1=0.38) methods, while post-hoc negation filtering paradoxically degraded LLM performance (Gemini+Hybrid Macro F1=0.58; Claude+Hybrid Macro F1=0.54).

Source-provided article image: Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes
medRxiv · Page 1

Interpretation

Within the same patient population, the same 46-symptom inventory, and the same expert-adjudicated gold standard, zero-shot LLM symptom extraction substantially outperformed dictionary-based rule matching and pretrained clinical NER. Prior work largely reported the feasibility of one method family at a time; this study provides a direct head-to-head benchmark of rule-based, NER, and LLM approaches on a common gold standard, using 2,704 notes as the cohort and 200 adjudicated notes as the evaluation set. Based on a 200-note gold standard independently annotated by two raters and jointly adjudicated (9,200 symptom-note pairs; pooled kappa=0.71, macro kappa=0.49), evaluated with Macro/Micro F1, precision, and recall, with pairwise comparisons via McNemar's test and Bonferroni correction (adjusted alpha=0.003).

Adding fixed-window rule-based negation filtering on top of LLM output reduced performance, driven mainly by a large increase in false negatives. The hybrid pipeline was designed as an exploratory test of whether rule-based negation filtering could further improve LLM precision beyond what contextual prompting already achieves; instead it overrode otherwise correct LLM predictions. Confusion matrices show false negatives rising from 53 to 163 for Claude and from 54 to 167 for Gemini, alongside reduced false positives (Gemini+Hybrid FP=75); the text illustrates this with examples such as 'no relief from nausea' and 'patient denies any improvement in abdominal pain,' where a fixed 100-character window cannot distinguish true negation from positive mentions that incidentally contain negation words.

Extraction difficulty varies markedly across symptoms, and the boundary between the general 'pain' item and anatomically site-specific pain items is hard to apply consistently in free text. The study disaggregates aggregate metrics to individual symptoms, noting that general pain (18.0% gold-standard prevalence) reached F1 of only 0.47 for Claude and 0.50 for Gemini, whereas abdominal_pain (0.94/0.98) and anal_rectal_pain (0.86/0.96) both exceeded 0.85. The same pattern appeared between human annotators: kappa for pain was 0.52 (moderate), lower than abdominal_pain (0.79) and anal_rectal_pain (0.86), suggesting an inherent ambiguity in the symptom taxonomy rather than a failure specific to any one method.

Gold-standard reliability itself varies with symptom prevalence, so evaluation of rare symptoms warrants careful interpretation. The study reports both pooled kappa (0.71) and macro-averaged kappa (0.49), and explains that six zero-prevalence symptoms necessarily yield F1=0 for every method, mechanically lowering macro-averaged F1. Only 277 of 9,200 symptom-note pairs (3.0%) were discordant; 20 symptoms had kappa>=0.60, while 12 had kappa<0.20 despite raw agreement of 94-99%, consistent with the kappa paradox under low prevalence; five symptoms had no positive ratings from either annotator, making kappa not estimable.

Perspective

The results are aimed at researchers and clinical informatics teams using routinely collected discharge notes for oncology symptom surveillance, in settings where the symptom inventory is the 46 items combined from the MSAS and EORTC QLQ-CR29 and where symptom information is documented mainly in the Chief Complaint and History of Present Illness sections; the methods run in a zero-shot setting without institution-specific rule development or model retraining, making them straightforward to reproduce and extend in environments with the appropriate data use agreements.

A careful reader would still watch how macro-averaged F1 changes on a larger evaluation set with better representation of rare symptoms; whether the boundary between general pain and site-specific pain items can be applied consistently through clearer annotation rules or prompt design; how proprietary model version churn (the text notes Gemini 2.5 Flash was retired during the study and superseded by Gemini 3.5 Flash) and transmission of clinical text to cloud APIs affect reproducibility and data governance; and whether locally deployable open-weight models could offer a more durable and privacy-preserving alternative. In addition, this is a preprint that has not been peer reviewed, and the specific values in Figure 1 and Figure 2 are not enumerated in the text, so symptom-level detail can only be inferred from the main text and supplementary tables.

Sources