Skip to main content
Back to timeline
medRxivSource publication:

Diagnostic Value of Large Language Model-Extracted Gross Brain Findings in Neurodegenerative Diseases

Synopsis

Using 5,613 autopsy cases from the Mayo Clinic Brain Bank collected between 1998 and 2023, this study fine-tuned a large language model to convert narrative gross descriptions into semi-quantitative scores for 39 features (extraction accuracy 0.95 on 200 manually annotated feature-level test examples), then classified seven neuropathologic diagnostic categories with a CatBoost classifier and a second fine-tuned LLM, both including age at death, sex, and brain weight; on a held-out test set of 562 cases CatBoost reached accuracy 0.73, kappa 0.65, and macro-average AUC 0.92, while the text-based LLM reached accuracy 0.75 and kappa 0.68, with macro-average sensitivity 0.66 for both, PSP sensitivity 0.92 and 0.93, MSA sensitivity 0.87 and 0.92, but AD-LBD sensitivity only 0.21 and 0.

Source-provided article image: Diagnostic Value of Large Language Model-Extracted Gross Brain Findings in Neurodegenerative Diseases
Fig. 1 ·

Fig. 1 Study cohort selection and model development. A total of 7,379 Mayo Clinic Brain Bank

medRxiv · Page 28

Interpretation

Narrative gross brain descriptions can be converted into structured features and used for neuropathologic diagnostic classification. The diagnostic value of gross findings across major neurodegenerative diseases had been incompletely characterized; this work systematizes free text from autopsy reports into semi-quantitative scores for 39 features. Based on 5,613 Mayo Clinic Brain Bank autopsy cases (1998–2023), with extraction accuracy of 0.95 on 200 manually annotated feature-level test examples.

Two classification routes, one using extracted scores and one using standardized text, performed similarly on the held-out test set. It compares a CatBoost numeric classifier with a second fine-tuned LLM text classifier, both including age at death, sex, and brain weight. On the 562-case held-out test set, CatBoost achieved accuracy 0.73, kappa 0.65, and macro-average AUC 0.92; the text-based LLM achieved accuracy 0.75 and kappa 0.68; macro-average sensitivity was 0.66 for both.

Diagnostic performance varied markedly across diseases, with better identification of disorders having distinctive macroscopic patterns. It quantifies per-category sensitivity differences across the seven diagnoses rather than reporting only aggregate metrics. PSP sensitivity was 0.92 (CatBoost) and 0.93 (text-based LLM), MSA was 0.87 and 0.92, whereas AD-LBD was only 0.21 and 0.03.

Feature attribution linked specific gross features to specific disease predictions. Feature attribution analysis identified subthalamic nucleus atrophy and putaminal abnormalities as contributors to PSP and MSA predictions, respectively. Attribution results come from the above models on the held-out test set and represent model-level associations rather than independent validation of pathologic mechanisms.

Perspective

This work targets gross description text in autopsy reports and applies to classification support for diseases with distinctive macroscopic patterns (such as PSP and MSA); its setting is a retrospective analysis of brain bank autopsy cases with age at death, sex, and brain weight included as additional variables, and for the combined AD-LBD category this route has limited discriminative ability.

The currently available text is at the abstract level and lacks figures and full methodological detail, so the exact computation of feature attribution, the per-category sample size distribution, and model performance on external data remain open questions; in addition, the very low AD-LBD sensitivity (0.21 and 0.03) suggests that the distinguishability of this category warrants further attention.

Sources