Skip to main content
Back to timeline
medRxivSource publication:

Privacy-Aware Distillation of Large Language Models for Enhanced Multimorbidity Scoring

Synopsis

This study introduces and evaluates a privacy-preserving knowledge distillation framework in which CTGAN-generated synthetic cohorts matching UK Biobank distributions are used to elicit multimorbidity scores from three teacher LLMs (GPT-4o, Gemini, DeepSeek) under zero-shot prompting, and compact student models (CoLLMs) are then trained to mimic those scores, enabling application to real UK Biobank data (N = 439,221) for multimorbidity scoring without exposing patient-level data to third-party APIs, with evaluation against the Charlson (CCI) and Elixhauser (ECI) indices via survival analysis, genome-wide association studies, and polygenic risk score associations.

Source-provided article image: Privacy-Aware Distillation of Large Language Models for Enhanced Multimorbidity Scoring
Figure 1 ·

Figure 1: Overview of the CoLLM framework for multimorbidity risk prediction and

medRxiv · Page 7

Interpretation

The study builds a CoLLM distillation pipeline: ICD-10 codes are grouped into 263 parent disease categories, CTGAN generates 50,000 training and 10,000 test synthetic records, teacher models produce scores under zero-shot prompts, and student models are multilayer perceptron regression networks trained to mimic each individual teacher and the mean of all three (LLM-Mean). Compared with directly feeding patient data to LLMs or relying only on rule-based scoring, this work transfers scoring capability from closed-source teacher models into a compact, locally deployable student model, thereby working around UK Biobank data-use restrictions on third-party APIs. Synthetic data fidelity was assessed with the SDV framework: overall quality score 95.34%, Column Shapes Score 96.69%, Column Pair Score 93.99%; disease prevalence across 263 parent categories correlated with the real cohort (Pearson r = 0.73, p < 0.001), and pairwise disease co-occurrence correlation was moderate (Pearson r = 0.58, p = 1.50e-03).

On held-out test data, CoLLMs reproduced teacher scores with strong agreement, with LLM-Mean performing best. This quantifies the fidelity ceiling of output-level distillation and indicates that ensembling multiple teachers yields more stable behavior than mimicking any single teacher. Spearman correlations ranged from 0.75 (CoLLM-DeepSeek) to 0.89 (CoLLM-Mean), with CoLLM-Gemini at 0.80 and CoLLM-GPT-4o at 0.77; R² values were 0.69 for LLM-Mean, 0.50 for Gemini, 0.42 for GPT-4o, and 0.24 for DeepSeek; CoLLM-Mean residuals showed near-zero correlation (ρ = 0.055), while individual CoLLMs, especially DeepSeek, progressively underpredicted higher multimorbidity scores.

On real UK Biobank data, CoLLM-derived scores improved all-cause mortality discrimination over CCI and ECI and showed interpretable genetic associations. The work links LLM-derived scores to three external lines of evidence—survival outcomes, genetic architecture, and polygenic risk scores—rather than reporting only correlation with teacher models. Median follow-up was 4.8 years (IQR 2.7–9.2) with 37,094 death events; in the 50–60 year group, LLM-Mean reached a C-index of 0.910 (95% CI 0.906–0.915) versus CCI 0.895 (0.889–0.900) and ECI 0.877 (0.870–0.883); GWAS showed consistent signals at HLA-DQA1/DQB1 and LPA; SNP heritability estimates were 0.034 ± 0.0023 for CCI, 0.046 ± 0.0026 for DeepSeek, 0.041 ± 0.0023 for Gemini, 0.049 ± 0.0025 for GPT-4o, and 0.048 ± 0.0025 for LLM-Mean; LLM-Mean identified 30 significant PRS associations after Bonferroni correction, more than CCI (22) and ECI (14).

An independent LLM-as-Judge evaluation found teacher LLM outputs clinically meaningful but variable, while distilled student models were more stable in clinical alignment, score reliability, and outlier/safety risk. This adds qualitative, multi-dimensional evidence beyond numerical correlation and suggests student models can be more stable than their teachers. Claude Sonnet-4 served as an independent judge on 10,000 synthetic test samples using a fixed-batch test-retest design (10 batches of 100 records, each rated three times, giving 30 repeated evaluations per model per dimension); Krippendorff's alpha was 0.762 for clinical alignment, 0.927 for score reliability, and 0.859 for safety/outlier risk; CoLLM-DeepSeek improved over Teacher DeepSeek in clinical alignment (2.00 to 3.03), reliability (1.93 to 3.47), and safety/outlier risk (2.57 to 4.00).

Perspective

The framework targets clinical and epidemiological research settings that need to leverage closed-source LLM capabilities under privacy and data-use constraints, particularly institutions with large electronic health records that cannot send individual-level data to third-party APIs; its design goal is to let researchers deploy a compact student model locally to generate multimorbidity scores from ICD-10 codes and demographic variables for survival risk stratification, genetic association, and polygenic risk score analyses. The authors also state that teacher models only saw CTGAN-generated synthetic profiles and that performance on real UK Biobank data was inferred indirectly through CoLLM, so results should be read as evidence of feasibility and potential utility within the current study setting rather than proof of broad applicability across all populations, datasets, or LLM systems.

Several open questions remain for a careful reader: the synthetic cohort over-represents hypertension (73.1% synthetic vs 34.4% real) and falls while under-representing general symptoms and signs, arthrosis, ischaemic heart disease, and diabetes mellitus, and how this distributional skew affects teacher scoring and student distillation is not yet fully characterized; CTGAN was not trained with differential privacy and no formal privacy audits such as membership-inference or attribute-inference attacks were performed, so disclosure-protection scores (0.77–1.00 against a 0.50 baseline) should be treated as a heuristic risk assessment; CoLLM approximates scalar scores rather than full reasoning chains, making it output-level distillation; teacher correlation and LLM-as-Judge scoring primarily measure agreement and consistency rather than clinical correctness and may inherit biases from teacher or judge models; and GWAS was restricted to European-ancestry individuals with no cross-population fairness analyses across ethnicity, socioeconomic status, or ancestry, leaving external validation in diverse cohorts and clinician adjudication as future work.

Sources