Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

medRxiv

Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes

Using 2,704 colorectal cancer discharge notes from MIMIC-IV and a 46-symptom inventory derived from the MSAS and EORTC QLQ-CR29, this study benchmarked dictionary-based rule matching, pretrained clinical NER, zero-shot Claude Haiku and Gemini 3.5 Flash, and two hybrid variants (LLM output plus post-hoc rule-based negation filtering) against a 200-note two-rater adjudicated gold standard, finding that Gemini 3.5 Flash performed best (Macro F1=0.70, Micro F1=0.86, Macro Precision=0.74), followed by Claude Haiku (Macro F1=0.63, Macro Recall=0.71), both substantially outperforming rule-based (Macro F1=0.44) and NER (Macro F1=0.38) methods, while post-hoc negation filtering paradoxically degraded LLM performance (Gemini+Hybrid Macro F1=0.58; Claude+Hybrid Macro F1=0.54).