Integrating Metabolomics, Mendelian Randomization, and Machine Learning, a Study Flags Phenylalanine and Its Transporter SLC6A14 as Candidate Diagnostic and Therapeutic Targets in Pancreatic Cancer
Synopsis
Integrating plasma metabolomics, Mendelian randomization, and machine learning, this study identified phenylalanine as causally associated with pancreatic cancer among 55 plasma metabolites (IVW OR = 1.641, 95% CI 1.052–2.562, p = 0.029), derived eight related differentially expressed genes, built a random forest diagnostic model from 113 combinations of 12 algorithms (training AUC 0.994; validation AUCs 0.918, 0.983, 0.923), used SHAP to rank SLC6A14 as the top feature, and combined single-cell sequencing, simulated gene knockout, molecular docking, and molecular dynamics to suggest genistein binds SLC6A14 stably, with RT-qPCR confirming high expression of the five model genes in a BxPC-3 versus HPDE6-C7 cell pair.
Research design flowchart.
PubMedInterpretation
The study establishes phenylalanine as a plasma metabolite with genetic causal evidence for pancreatic cancer: inverse-variance weighted analysis using 15 SNPs as instrumental variables gave OR = 1.641 (95% CI 1.052–2.562, p = 0.029), with directionally consistent results from MR-Egger, weighted median, simple mode, and weighted mode; Cochran's Q showed no significant heterogeneity (IVW Q = 6.764, df = 14, p = 0.943), the MR-Egger intercept showed no directional horizontal pleiotropy (intercept = −0.044, SE = 0.051, p = 0.408), and leave-one-out analysis showed no single SNP dominating the estimate. Prior work on metabolites and pancreatic cancer has largely rested on observational differential analysis; this study chains metabolomic differential screening to Mendelian randomization causal inference, using genetic instruments to prioritize metabolites. Two-sample MR based on OpenGWAS dataset ieu-a-822 (3835 participants; 1896 cases and 1939 controls) with multiple complementary estimators and sensitivity analyses; causal strength is bounded by the p < 1 × 10⁻⁵ instrument threshold and the source population.
Using 12 machine learning algorithms across 113 algorithm combinations, the study built a pancreatic cancer diagnostic model; random forest was optimal, with training AUC 0.994 and validation AUCs of 0.918, 0.983, and 0.923, training sensitivity 94.7%, specificity 94.0%, accuracy 94.4%, and validation sensitivities of 87.2%, 95.6%, 84.6%, specificities of 87.2%, 88.9%, 100.0%, and accuracies of 87.2%, 92.2%, 90.5%. The model was built on candidate genes intersected from phenylalanine-related differentially expressed genes and the WGCNA red module (correlation with PC r = 0.72, p = 1e-80), giving feature selection a metabolic causal lead rather than a purely data-driven one. Training cohorts GSE16515, GSE62452, GSE183795 and validation cohorts GSE15471, GSE28735, GSE71989, with 10-fold cross-validation and Z-score standardization; the authors note the small GSE71989 validation cohort (n = 21) and a 0.011–0.076 gap between training and validation AUC.
SHAP interpretation ranked SLC6A14 as the largest contributor to model predictions, followed by SLC2A1 and GJB2; single-cell sequencing (GSE197177; one adjacent normal tissue and three primary untreated pancreatic ductal adenocarcinoma samples) showed SLC6A14 predominantly expressed in epithelial cells, and simulated knockout of SLC6A14 significantly affected 95 genes (1.9% of the gene pool), enriched in tight junctions, cell adhesion molecule interactions, and leukocyte transendothelial migration. The study converts model feature importance into an interpretable biological hypothesis, linking SLC6A14 to amino acid transport and immune cell communication and proposing it as a metabolic checkpoint connecting amino acid availability to immune evasion. Computational evidence from SHAP (TreeSHAP with a background set of 100 randomly sampled training instances), single-cell transcriptomics, and scTenifoldKnk virtual knockout; the authors state the mechanism still requires functional transport assays, targeted metabolomics, and in vivo validation.
Drug prediction retrieved 110 drug molecules from DSigDB and, after comparison with a traditional Chinese medicine active-compound database, yielded 19 candidate herbal extracts; molecular docking showed genistein and withaferin A bound SLC6A14 with affinities of 9.1 and 8.9 kcal/mol, and a 100 ns molecular dynamics simulation showed a stable complex (Rg 2.80 ± 0.03 nm, RMSD 0.91 ± 0.10 nm, RMSF 0.18 ± 0.17 nm, SASA 286.91 ± 6.84 nm²) with an MM/GBSA binding free energy of −33.77 ± 1.08 kcal/mol. The study pushes the core gene from the diagnostic model toward a candidate therapeutic target and lead compound, proposing the isoflavone/flavonoid scaffold as a privileged chemotype for SLC6A14 targeting. Entirely in silico evidence (GROMACS 5.0.4, CHARMM36m force field, TIP3P water model, 100 ns NPT simulation); the authors note genistein–SLC6A14 binding still requires experimental confirmation.
Perspective
The results are aimed at research settings in pancreatic cancer early molecular diagnosis and targeted therapy: the five-gene random forest model applies to risk stratification and diagnostic assessment based on GEO transcriptomic data, while the SLC6A14 findings apply to mechanistic studies and lead-compound screening centered on amino acid metabolism and the tumor immune microenvironment. The authors' proposed next steps include functional validation of SLC6A14 in pancreatic cancer cell lines, targeted metabolomic profiling and coculture assays to test the proposed metabolic–immune crosstalk, in vivo evaluation of genistein in pancreatic cancer xenograft models, and prospective clinical cohort validation of the five-gene model.
Readers should still watch the gap between training AUC 0.994 and validation AUCs of 0.918–0.983, and the effect of the small GSE71989 validation cohort (n = 21) on performance estimates; functional validation of SLC6A14-mediated phenylalanine metabolic reprogramming is still lacking, and genistein–SLC6A14 binding plus the hypothesis that elevated SLC6A14 depletes extracellular phenylalanine and impairs antitumor immunity rest on computational and association analyses that need functional transport assays, targeted metabolomics, and in vivo models; RT-qPCR validation was performed in only one cell pair (BxPC-3 versus HPDE6-C7). In addition, although this is a full-text parse, some result details (such as the relationship between the SHAP direction descriptions and gene expression direction) depend on figures and supplementary data, so precise reproduction should be checked against Data S2–S6 and Figures 9 and 14.
