Identifying cohorts at elevated risk of cancers using generative modeling of patient health states
Synopsis
This study introduces GenEHR, an autoregressive generative model trained on electronic health records from millions of patients that explicitly represents irregular inter-visit time intervals via RAdix Time Encoding and combines parameter-efficient gated low-rank adaptation for supervised fine-tuning, improving prediction of a first cancer diagnosis within a five-year horizon across five large EHR cohorts and supporting risk-based screening for aggressive cancers such as pancreatic and ovarian cancer.
Interpretation
GenEHR uses RAdix Time Encoding (RATE) to represent irregular inter-visit time gaps at day-level resolution using a small set of digit tokens, avoiding a separate token for every possible interval or artificial no-event tokens. Prior methods such as ETHOS and CoMET use a fixed vocabulary of time-interval tokens and cannot flexibly capture fine-grained gaps; Delphi relies on artificial no-event tokens and its performance depends on the rate at which these tokens are inserted. RATE covers 0 to 4,199 days using 27 output logits, balancing resolution and vocabulary size. Across five cohorts, event prevalence in generated trajectories reproduced observed prevalence with log-Pearson correlations of 0.93 to 0.99 and Spearman correlations of 0.90 to 0.98; generated inter-visit time gaps correlated with observed gaps with Spearman correlations of 0.89 to 0.98; time-encoding ablation showed RATE performed best for predicting the next event in the next visit, with the largest improvement at longer time gaps.
Jointly training a pan-cancer risk head and cancer-type-specific gated low-rank adapters (gated LoRA) over a frozen GenEHR backbone achieved better cancer risk discrimination than direct scoring and head-only adaptation across five cohorts and multiple prediction horizons. Direct scoring using frozen GenEHR cancer-token logits can discriminate cases from controls without cancer-registry supervision; adding time-to-event supervision substantially improved performance, showing that explicit supervision on diagnosis timing adds information. Gated LoRA combines shared low-rank directions with cancer-specific gates, allowing rare cancers to benefit from shared representations supported by other cancers. In MGB, macro-AUROCs across horizons ranged from 0.71 to 0.77 for direct scoring, 0.80 to 0.84 for head only, and 0.80 to 0.84 for gated LoRA; PROV ranged from 0.60 to 0.67, 0.72 to 0.77, and 0.73 to 0.78; VA ranged from 0.67 to 0.71, 0.80 to 0.81, and 0.82 to 0.83; AoU ranged from 0.58 to 0.85, 0.67 to 0.82, and 0.68 to 0.84; UKBB ranged from 0.54 to 0.80, 0.60 to 0.92, and 0.62 to 0.92.
At a clinically relevant high-risk operating threshold, GenEHR-CancerRisk substantially enriches cancer cases among the top 1,000 highest-risk patients and provides lead-time distributions from prediction cutoff to actual diagnosis. Unlike reporting AUROC alone, the study reports SIR, SNS, and PPV at an operating threshold of N=1000, directly addressing the need for screening programs to choose an operating threshold that trades false positive rate against true positive rate and accounts for healthcare system capacity. At the 60-month prediction horizon and N=1000 threshold, SIR and SNS ranged from 18-224 and 2-96 in MGB, 6-224 and 2-136 in PROV, 3-91 and 6-299 in VA, 3-27 and 5-570 in UKBB, and 4-20 and 6-322 in AoU; for pancreatic, ovarian, and lung cancer, SIR ranged from 34-48, 6-22, and 20-57, while SNS ranged from 13-18, 36-136, and 2-7; median lead time ranged from 5.5 to 29.8 months and maximum lead time ranged from 21.4 to 60 months.
GenEHR's learned patient and event representations organize by clinical category without explicit clinical hierarchy supervision, and a conditional pointwise mutual information belief graph can generate hypotheses about events preceding cancer diagnosis. Event embeddings in UMAP projection clustered by ICD chapter, medication ATC class, lab assay group, and CPT procedure section, with refined substructure separating acute from chronic respiratory disease, sprains from fracture, and upper gastrointestinal symptoms from hernia and hepatobiliary disease; the cPMI belief graph showed pancreatic and biliary tract disorders as top associations to pancreatic cancer, a brain cancer neighborhood including anti-seizure medicines and intracranial pressure treatment, and a myeloid leukemia neighborhood including aplastic anemia and bone-marrow failure. cPMI edges represent learned model belief about the cohort and do not imply causality; some edges agree with prior clinical knowledge, while others may generate hypotheses for further epidemiological or experimental investigation.
Perspective
This is a retrospective analysis evaluated on held-out test sets from five cohorts; the results support evaluation of GenEHR-CancerRisk as a prospective clinical decision-support tool for prioritizing patients for risk-based screening for aggressive cancer types such as pancreatic and ovarian cancer. The model applies to patient populations with sufficient longitudinal history in routinely collected structured EHR data and requires local training and calibration within healthcare systems that have cancer registry-confirmed outcomes.
This preprint has not been peer reviewed, and retrospective results cannot be directly translated into clinical practice; model performance varies substantially across cancer types and cohorts, with wide SNS ranges for some cancers; cPMI belief graph edges reflect temporal associations learned by the model rather than causal relationships; the accuracy of EHR diagnosis codes remains challenging in U.S. healthcare systems; the model did not use clinical notes or medical images; future prospective silent evaluation is needed to assess reliability, temporal drift, calibration, and performance across demographic, socioeconomic, and environmental subpopulations.
