Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

medRxiv

Diagnostic Value of Large Language Model-Extracted Gross Brain Findings in Neurodegenerative Diseases

Using 5,613 autopsy cases from the Mayo Clinic Brain Bank collected between 1998 and 2023, this study fine-tuned a large language model to convert narrative gross descriptions into semi-quantitative scores for 39 features (extraction accuracy 0.95 on 200 manually annotated feature-level test examples), then classified seven neuropathologic diagnostic categories with a CatBoost classifier and a second fine-tuned LLM, both including age at death, sex, and brain weight; on a held-out test set of 562 cases CatBoost reached accuracy 0.73, kappa 0.65, and macro-average AUC 0.92, while the text-based LLM reached accuracy 0.75 and kappa 0.68, with macro-average sensitivity 0.66 for both, PSP sensitivity 0.92 and 0.93, MSA sensitivity 0.87 and 0.92, but AD-LBD sensitivity only 0.21 and 0.