Skip to main content
Back to timeline
Archives of toxicologySource publication:

Artificial Intelligence in Toxicology: Current Advances, Challenges and Future Directions

Synopsis

This review, based on a structured literature search of PubMed/MEDLINE, Web of Science and Scopus (2000-2026) supplemented by guidance documents from OECD, FDA, EMA, EFSA and EPA, traces the development of AI in toxicology from rule-based expert systems to deep learning, large language model and multimodal architectures, reports that graph neural networks and multi-task deep learning have shown competitive performance in selected benchmark studies of drug-induced liver injury, hERG cardiotoxicity and Ames mutagenicity, and concludes that inconsistent external validation, limited generalisation of endpoint-specific models and hallucination in large language models leave unresolved regulatory risks, so AI models augment but cannot yet replace experimental toxicology.

AI-generated editorial illustration: Artificial intelligence in toxicology: current advances, challenges and future directions.

Interpretation

The review systematically traces the evolution of AI in toxicology from rule-based expert systems to contemporary deep learning, large language model and multimodal architectures, and evaluates the evidence base for AI-based toxicity prediction. Compared with reviews focused on a single method or endpoint, it combines a structured literature search (PubMed/MEDLINE, Web of Science, Scopus, 2000-2026) with regulatory guidance documents, placing methodological evolution and regulatory relevance in one framework. A narrative review whose evidence comes from a structured literature search and regulatory guidance documents rather than new experimental or benchmark data.

Graph neural networks, multi-task deep learning and related AI approaches have shown competitive performance in selected benchmark studies of drug-induced liver injury (DILI), hERG cardiotoxicity and Ames mutagenicity. It locates the competitiveness of AI toxicity prediction at the level of benchmark studies for specific endpoints (DILI, hERG, Ames) rather than claiming broad superiority over traditional methods. Based on reported performance in the selected benchmark studies, while the text explicitly notes that performance varies substantially across datasets and validation settings.

Explainable AI frameworks such as SHAP are aligning model outputs with adverse outcome pathways (AOPs), and federated learning, as demonstrated by the MELLODDY consortium, enables privacy-preserving multi-institutional collaboration. It presents explainability and federated learning as routes connecting AI predictions to toxicological mechanistic frameworks and to cross-institutional data collaboration. Illustrative evidence from SHAP and the MELLODDY consortium, presented as a direction rather than a systematic evaluation of effects.

Inconsistently reported external validation, difficulty generalising endpoint-specific models and hallucination in large language models constitute unresolved regulatory risks, and realising the potential requires adherence to the TREAT validation principles and sustained regulatory engagement. It explicitly frames validation norms and regulatory engagement as preconditions for routine application, and positions AI as augmenting rather than replacing experimental toxicology. A review-level judgement based on critical evaluation of the existing evidence base and regulatory guidance documents.

Perspective

The review addresses researchers, model developers and regulatory participants in toxicological safety assessment, and applies to the setting of AI prediction of safety-relevant signals for drugs, chemicals and environmental contaminants; its conclusions rest on a structured literature search and regulatory guidance documents, serving to clarify methodological evolution and differences in readiness rather than to provide directly applicable operating standards.

The available text is at the abstract level and does not detail specific datasets, sample sizes, model architectures or endpoint performance values, so the actual performance differences across DILI, hERG and Ames endpoints cannot be quantified here; the precise scope of inconsistent external validation and of generalisation difficulties for endpoint-specific models, the practical regulatory impact of large language model hallucination, and how the TREAT validation principles will be implemented remain open questions to watch.

Sources