Skip to main content
Back to timeline
Cancer DiscoverySource publication:

MutationProjector, pretrained on 30,000+ tumor genomes, predicts immunotherapy response in bladder, lung and melanoma cancers and flags KMT2D and SMARCA4-STK11 as candidate biomarkers

Synopsis

The authors built MutationProjector, a graph-attention cancer-genome foundation model integrating eight molecular network types, pretrained by masked gene reconstruction on genetic alterations from 30,328 tumors across 10 solid cancer types, then transfer-learned with a random forest on small labeled cohorts; it stratified survival in immunotherapy response prediction with hazard ratios of 0.71 for bladder, 0.74 for lung and 0.59 for melanoma, reached HR = 0.43 in a cisplatin-treated bladder cohort, and proposed KMT2D and SMARCA4-STK11 as candidate biomarkers via attention analysis.

Source-provided article image: A Foundation Model of Cancer Genotype Enables Precise Predictions of Therapeutic Response.

(A, B) Datasets collected to pre-train MutationProjector, including (A) tumor genetic alterations and (B) molecular networks. CNA: Copy Number Amplifications, CND: Copy Number Deletions, DDRAM ( 65 ): DNA Damage Response Assemblies Map, ISLE ( 63 ): Identification of clinically relevant Synthetic LEthality, PCNet ( 23 ): Parsimonious Composite Network, SIGNOR ( 58 ): SIGnaling Network Open Resource, STRING ( 66 ): Search Tool for Retrieval of Interacting Genes/Proteins, TMB: Tumor Mutation Burden, TRRUST ( 59 ): Transcriptional Regulatory Relationships Unravelled by Sentence-based Text-mining. (C) Model configuration for pre-training. Solid black mutation embedding indicates alteration profiles that are masked. Updated embeddings are generated via network-based message passing using graph attention networks. Size of the gene and covariate embeddings are noted as d features . CNA: Copy Number Amplification, CND: Copy Number Deletion, Emb: Embedding, Norm: Normalization, TIL: Tumor Infiltrating Lymphocytes, TMB: Tumor Mutational Burden.

PubMed

Interpretation

MutationProjector maps a tumor's alteration status across 468 clinical-panel genes plus TMB, aneuploidy and mutational signatures into an interpretable quantitative embedding via a graph attention network, trained with masked gene reconstruction as a self-supervised objective. Prior tumor-genome models were typically trained for a single task or data type; this work unifies multiple molecular network types (physical interaction, transcriptional regulation, phosphorylation, ubiquitination, genetic interaction, DNA damage repair, STRING, PCNet, totaling 19,789 interactions) with genomic covariates in one pretraining framework. Pretraining used 30,328 tumors split 80%/20% into training and held-out sets (24,262/6,066); masked gene mutation prediction reached AUPRC 0.21 versus 0.02 for random guessing, a 9.7-fold improvement, with similar performance in an external cohort.

Without being given the relevant labels, the pretrained embedding spontaneously separates tissue types, squamous clusters, HPV status, and basal versus luminal transcriptional subtypes in bladder and breast cancer, while retaining key driver alteration information. These structures are not driven by explicit supervision: the embedding still stratifies major cancer types when the cancer-type prediction task is removed, and HPV status and mRNA-defined subtypes were never model inputs. Based on UMAP projections of the embedding, with MSK-IMPACT and TCGA samples clustering closely in latent space, suggesting minimal batch effects; the embedding contributed to predicting mutation patterns beyond tissue type for 95.7% of genes.

After fine-tuning on small labeled cohorts, the same pretrained model matches or exceeds current biomarkers and classical machine learning methods across multiple independent cohorts for immunotherapy, chemotherapy and metastasis-related tasks. It follows a 1-model/N-tasks paradigm rather than building separate models per task; the immunotherapy classifier trained on only 94 patients generalized to three independent bladder, lung and melanoma cohorts. Immunotherapy hazard ratios were 0.71 for bladder (p=1.6×10⁻²), 0.74 for lung (p=1.7×10⁻⁵) and 0.59 for melanoma (p=5.8×10⁻³); chemotherapy bladder HR=0.43 (p=2.5×10⁻²); metastasis prediction reached AUPRC 0.84 in an independent cohort; clinical tasks totaled 3,362 patients.

Attention analysis yields an interpretable biomarker list, including chromatin remodelers such as KMT2D and SMARCA4 linked to immunotherapy sensitivity, and co-alteration combinations KRAS-STK11, KEAP1-STK11 and SMARCA4-STK11 predicting non-response. KRAS, STK11 or KEAP1 mutations were not individually more frequent in non-responders, but their pairwise co-alterations were; 94% of high-attention gene pairs would not have been prioritized by standard co-mutation analysis. Feature importance was assessed via attention-weight-based linear probing and Spearman correlation; KRAS-STK11, KEAP1-STK11 and SMARCA4-STK11 accounted for 14%, 13% and 10% of predicted non-responders, respectively.

Perspective

The results apply to solid tumors profiled with clinical gene panels; pretraining covered 10 solid cancer types (centered on breast, colorectal and non-small-cell lung cancer), and downstream tasks focus on immunotherapy, cisplatin chemotherapy and metastatic outcomes. For cancer types not yet included, such as pancreatic cancer, prostate cancer or sarcomas, and for new settings such as liquid biopsy, the authors frame these as expansion directions. The intended audience is researchers and translational teams seeking to turn routine panel data into therapy-response predictions and biomarker hypotheses.

Clinical validation cohorts are limited in size (only 94 patients for immunotherapy training and 42 for chemotherapy testing), so the hazard ratios and p-values rest on these cohorts and robustness of extrapolation still needs larger prospective data; biomarkers such as KMT2D and SMARCA4-STK11 come from model attention analysis and are hypothesis-generating rather than mechanistically confirmed; details of embedding dimension, attention aggregation and feature-importance methods depend on supplementary figures, so readers of the main text alone may need to consult the supplement to verify some quantitative bases.

Sources