A Graph-based QSAR Modeling Pipeline for Predicting In vitro PubChem Assays and In vivo Human Hepatotoxicity: Mechanistic Analysis of Caspase-3/7 Activation
Synopsis
This study developed a graph-based QSAR modeling pipeline integrating assay data preprocessing, fingerprint and molecular graph feature representations, and benchmarking of classical machine learning, graph neural networks, graph transformers, and their consensus ensembles, applied to predict Caspase-3/7 activation, mitochondrial membrane potential disruption, and FDA drug-induced liver injury, where Graphormer achieved the highest F1 of 0.79 and the full consensus model achieved the highest AUC of 0.69 on DILI prediction, surpassing the previous best model with AUC 0.63 and F1 0.65, and identified structural motifs associated with dual activation and cell-line-specific responses through fragment enrichment analysis.
Interpretation
Proposed and released an end-to-end graph-based QSAR modeling pipeline covering data cleaning, deduplication, feature representation, and benchmarking of multiple model classes. No prior QSAR study had comprehensively evaluated fingerprint-based and graph-based approaches (GNNs and GTs) for Caspase and mitochondrial toxicity prediction, filling this gap. The pipeline was applied across 19 qHTS assays, 18 final datasets, and 14 cell lines, with code and data publicly released; models included 3 classical classifiers, 4 GNNs, 2 GTs, and 4 consensus models.
On the FDA drug-induced liver injury benchmark, graph transformers and consensus ensembles outperformed the previous best model. Graphormer achieved the highest F1 of 0.79 and the full consensus model achieved the highest AUC of 0.69, compared to the previous best multi-feature DILIPredictor with AUC 0.63 and F1 0.65. Evaluated on a held-out test set of 222 compounds (155 hepatotoxic, 67 non-hepatotoxic) using a scaffold-based split to reduce train-test similarity; the Proxy+DILI setting consistently improved over Proxy-only.
Identified structural motifs associated with dual Caspase and mitochondrial activation through cross-assay mechanistic analysis. Prior work combined 13 cell lines into pan-Caspase toxicity without distinguishing cell-type differences; this work identified 16 significantly enriched fragments, including para-hydroxyphenyl (padj=5.0×10⁻⁷) and the lipophilic chain family. BRICS fragment enrichment with one-sided Fisher's exact test and Benjamini–Hochberg correction, contrasting 144 dual-active against 2,628 dual-inactive compounds; reverse analysis returned no significant hits.
Identified cell-line-specific structural motifs for Caspase-3/7 activation. HEK293 uniquely enriched 1,1-dichloroethane and chlorobenzene, SK-N-SH uniquely enriched an epoxide fragment, and H-4-II-E uniquely enriched a tetramethylcyclohexene motif and an acetaldehyde fragment, providing fine-grained evidence of cross-cell-line differences. BRICS enrichment applied across 8 cell lines with FDR set at α=0.10 to accommodate small samples (n=10–19); HepG2, CHO-K1, Jurkat, and SH-SY5Y returned zero significant fragments.
Perspective
The pipeline is suited to toxicity endpoints with public qHTS concentration-response data, performing strongly when the number of active compounds is relatively large (e.g., over 1,400 active compounds in the mitochondrial toxicity assay); performance is limited when active compounds are scarce and class imbalance is severe (e.g., Caspase-3/7 datasets). Applicability domain analysis shows that restricting predictions to within-domain compounds improves performance (ROC-AUC from 0.91 to 0.93), while out-of-domain compounds represent structurally novel scaffolds requiring experimental confirmation. The framework targets computational toxicology researchers and drug safety assessors for toxicity prioritization and mechanistic hypothesis generation.
F1 values on Caspase-3/7 datasets are low (0.08–0.21), and reducing false negatives under severe class imbalance remains an open question; Graphormer performed anomalously low on the HepG2 Caspase data without a clear explanation from the authors; some cell lines in the cell-line-specific motif analysis had small sample sizes (n=10–19), with the FDR threshold relaxed to α=0.10; in vitro to in vivo hepatotoxicity generalization is limited, and the authors propose future work combining literature mining and LLM approaches. Readers may watch whether these directions yield larger-scale data and external validation.
