Fully Automated Abstraction of Longitudinal Breast Oncology Records with Off-The-Shelf Large Language Models
Synopsis
The study developed a HIPAA-compliant open-source pipeline in which off-the-shelf commercial large language models, without fine-tuning, abstracted variables from unnormalized, unlabeled, and unedited clinical notes, pathology reports, medication administration records, and demographics for 100 complex breast cancer patients (median chart over 3,100 pages, median 6.5 years of follow-up, median 7 lines of therapy), achieving high concordance with an expert oncologist for recurrence status (99%), germline BRCA1/2 pathogenic variants (100%), hormone receptor status (99%), HER2 status (96%), clinical stage (91%), PIK3CA mutation status (91%), and ESR1 mutation status (90%), approaching inter-oncologist variability for anti-cancer drug extraction, while exact therapy-line reconstruction remaine
Interpretation
Within a fixed retrieval pipeline, off-the-shelf commercial LLMs can abstract a range of key variables from complex longitudinal oncology records with performance approaching inter-oncologist variability, without any fine-tuning or institution-specific retraining. Prior chart abstraction relied largely on manual review or models trained on institution-specific data; this work shows general off-the-shelf models directly processing unnormalized, unlabeled, and unedited raw text in a HIPAA-compliant setting. Based on 100 randomly selected patients from a complex-care-enriched breast cancer cohort, with breast oncologist abstraction as the reference standard and reported per-variable concordance percentages.
For systemic therapy extraction, all four tested LLMs outperformed research coordinators, and the best-performing LLM approached inter-oncologist variability. Model performance is benchmarked against both a second oncologist and research coordinators, rather than a single reference standard. Systemic therapy extraction included a second oncologist and research coordinators as comparators, with reported relative performance of the models versus coordinators.
LLM-derived datasets produced similar recurrence-free survival, overall survival, and hazard ratio estimates to expert-derived datasets. The evaluation moves from the variable level to downstream survival analysis, indicating that abstraction differences may not alter epidemiologic conclusions. Survival and hazard ratio estimates were compared between expert-derived and LLM-derived datasets in the same 100 patients.
In an external cohort of 97 young patients with early-stage breast cancer, the unmodified pipeline showed similar performance for recurrence detection and adjuvant endocrine therapy use. Provides initial out-of-institution validation, suggesting the pipeline does not depend on a single institution's chart format. External cohort of 97 patients, evaluated on recurrence detection and adjuvant endocrine therapy use.
Perspective
The results apply to HIPAA-compliant settings where a fixed retrieval pipeline processes complex longitudinal breast oncology records to extract variables such as diagnosis and recurrence dates, clinical stage, biomarker subtype, genetic testing results, and systemic therapies; external validation is limited to recurrence detection and adjuvant endocrine therapy use in 97 young patients with early-stage breast cancer. Exact therapy-line reconstruction remains a relatively lower-performing task, suggesting the pipeline is better suited to variable-level and survival-analysis-level dataset construction than to fully replacing line-by-line treatment curation.
The loaded text is a fast parse without figures or supplementary materials, so per-variable confidence intervals, model version details, prompt design, error-type distributions, and the full variable range of the external cohort cannot be checked; whether the 9-percentage-point gap in exact therapy-line reconstruction relative to the second oncologist is acceptable depends on the specific clinical or research use case; and the influence of different institutions' chart formats, languages, and coding conventions awaits further external validation.
