Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

medRxiv

Privacy-Aware Distillation of Large Language Models for Enhanced Multimorbidity Scoring

This study introduces and evaluates a privacy-preserving knowledge distillation framework in which CTGAN-generated synthetic cohorts matching UK Biobank distributions are used to elicit multimorbidity scores from three teacher LLMs (GPT-4o, Gemini, DeepSeek) under zero-shot prompting, and compact student models (CoLLMs) are then trained to mimic those scores, enabling application to real UK Biobank data (N = 439,221) for multimorbidity scoring without exposing patient-level data to third-party APIs, with evaluation against the Charlson (CCI) and Elixhauser (ECI) indices via survival analysis, genome-wide association studies, and polygenic risk score associations.