Skip to main content
Back to timeline
medRxivSource publication:

Large Language Model-derived Symptom Clusters and Patient Outcomes in Colorectal Cancer from MIMIC-IV Clinical Notes

Synopsis

Using a zero-shot large language model pipeline (Gemini 3.5 Flash and Claude Haiku) to extract 46 symptoms from 2,728 discharge notes of 1,507 colorectal cancer patients in MIMIC-IV, this study built patient-level symptom co-occurrence networks with phi correlation (≥0.10) and Louvain community detection; both models converged on three clinically coherent symptom clusters — Systemic, CRC Disease-Specific, and Gastrointestinal — and Systemic cluster burden was associated with in-hospital mortality (OR=1.33) and 1-year mortality (OR=1.41) while CRC Disease-Specific cluster burden independently predicted 30-day readmission (OR=1.20), with both associations robust to adjustment for metastatic disease.

AI-generated editorial illustration: Large Language Model-derived Symptom Clusters and Patient Outcomes in Colorectal Cancer from MIMIC-IV Clinical Notes

Interpretation

Symptoms extracted by large language models from unstructured discharge notes recover three clinically coherent symptom clusters: Systemic, CRC Disease-Specific, and Gastrointestinal. Prior colorectal cancer symptom cluster research relied mainly on patient-reported outcome surveys at discrete assessment points; this study instead used free-text electronic health record narratives generated during routine care, extracting symptoms with zero-shot large language models and then applying network analysis. Based on 1,507 patients and 2,728 discharge notes; the extraction method was previously benchmarked against a manually annotated 200-note gold standard (best model Macro F1=0.70); two independent large language models (Gemini 3.5 Flash and Claude Haiku) produced a consistent three-cluster structure with cross-model Adjusted Rand Index=0.727.

The three symptom clusters were reproducible across sensitivity analyses: mean ARI=0.985 across 100 Louvain random seeds, and 96.0% of 200 bootstrap resamples recovered the three-cluster solution (mean ARI=0.733), with the Gastrointestinal cluster identical in membership across both models. Few prior studies systematically tested the stability of networks built from large language model-extracted symptoms; this study prespecified and reported sensitivity analyses across correlation thresholds, random seeds, note-aggregation strategy, and bootstrap resampling. A three-cluster solution appeared at both phi≥0.10 and phi≥0.15 (though membership shifted, ARI=0.438), while at phi≥0.20 the network fragmented into five clusters; cross-model differences were confined to redistribution of a few boundary symptoms (sweats, feeling drowsy, dizziness, problems with urination).

Cluster burden carried independent prognostic value: Systemic cluster burden predicted in-hospital mortality (OR=1.33, 95% CI 1.18–1.49) and 1-year mortality (OR=1.41, 95% CI 1.27–1.56), and CRC Disease-Specific cluster burden independently predicted 30-day readmission (OR=1.20, 95% CI 1.08–1.34). Prior work often stopped at identifying and describing symptom clusters; this study entered all three cluster burden scores simultaneously into logistic regression to test predictive validity for three clinical outcomes. N=1,507, with outcome rates of 10.3% in-hospital mortality, 27.1% 30-day readmission, and 15.5% 1-year mortality; models included all three cluster burdens simultaneously and adjusted for age and sex, and the associations remained robust after additional adjustment for metastatic disease.

CRC Disease-Specific cluster burden was inversely associated with 1-year mortality (OR=0.85; OR=0.83 after adjustment for metastatic disease), and this inverse association was confined to patients with metastatic disease (OR=0.81), while metastatic disease was more, not less, prevalent among those with higher CRC Disease-Specific burden (43.5% at burden 0 rising to 78.8% at burden 3). The authors tested and reported the alternative explanation of confounding by disease stage, found the data did not support it, and proposed instead a hypothesis about the form of symptom expression (discrete, clinically actionable bowel symptoms versus generalized systemic decline). Metastatic disease was itself a strong predictor of 1-year mortality (OR=4.57, 95% CI 3.17–6.58); the inverse association strengthened rather than weakened after adjustment and remained evident within the metastatic subgroup, which the authors use to argue against a less-advanced-disease explanation.

Perspective

The results apply to colorectal cancer inpatients treated at a single academic medical center whose symptoms are documented in discharge summaries, with symptoms drawn from the Chief Complaint and History of Present Illness sections, extracted by zero-shot large language models, and organized into patient-level networks via phi correlation and Louvain community detection. This enables subsequent work to perform scalable symptom phenotyping on the same class of electronic health record text and to treat cluster burden scores as candidate risk markers for mortality and readmission; it is most directly relevant to clinical and nursing researchers who want to use automated symptom profiling for symptom assessment and patient-reported outcome monitoring. The authors note that an important next step is to move beyond cross-sectional co-occurrence toward longitudinal symptom dynamics, using sequential electronic health record notes to identify temporally antecedent or sentinel symptoms.

Symptoms were aggregated across all discharge notes without accounting for temporality, and under first-note restriction mean symptom burden fell from 4.15 to 2.73 with several symptoms dropping below the 5% prevalence threshold, so the temporal relationship between symptom occurrence and outcomes remains an open question. Psychological symptoms were not prominent in these large language model-derived networks, whereas prior patient-reported studies and one electronic health record study using MetaMap identified an anxiety–insomnia–weakness–depression cluster; the authors call for direct comparison of large language model-extracted symptoms with contemporaneous patient-reported outcomes to distinguish documentation bias from extraction error. The phi correlation (≥0.10) and prevalence (≥5%) thresholds were prespecified rather than empirically optimized, and clusters were derived with a variable-centered approach linked to outcomes via unweighted burden scores, which differs conceptually from person-centered phenotypes such as latent class analysis. In addition, this is a preprint that has not been peer reviewed, the findings come from a single academic medical center, and external validation across health systems has not yet been completed.

Sources