Skip to main content
Back to timeline
Frontiers in ToxicologySource publication:

Repeated ChatGPT runs for the ATBC oral PDE returned values from 5 to 300 mg/day, prompting a modular, human-supervised LLM workflow for toxicological risk assessment

Synopsis

A collaborative working group reviewed the principles, strengths, and limits of large language models (LLMs) in toxicological risk assessment (TRA) and then used an exemplar case study, the derivation of the oral Permitted Daily Exposure (PDE) for acetyl tributyl citrate (ATBC, CAS 77-90-7) under the draft ICH Q3E guideline, testing zero-shot end-to-end prompts, structured step-guided prompts, and a document-constrained configuration with ChatGPT (GPT-5 and later GPT-5.

Source-provided article image: A pragmatic workflow for integrating large language models into toxicological risk assessment
FIGURE 1

FIGURE 1 From case-study observations to a proposed LLM-powered workflow for TRA. The study foundations comprise a background analysis of LLMs in the context of TRA, the regulatory context, and the perspectives of the working group. An exemplar case study on the derivation of a Permitted Daily Exposure (PDE) for acetyl tributyl citrate was then conducted. Exploratory LLM investigations generated observations on the use of LLMs for core TRA tasks, including evidence retrieval, summarization, and reasoning. These observations were interpreted in the context of the available literature to identify both strengths and limitations of LLM-assisted approaches and to inform potential mitigation strategies. Together, the study foundations, case-study observations, and their interpretation informed the workflow hypothesis that LLMs may be most appropriately integrated into TRA as components of structured, modular, human-supervised workflows that support, rather than replace, expert judgment. Exploratory investigations were conducted primarily using ChatGPT, with selected Claude and Gemini runs reported in the Supplementary Material.

· Page 3

Interpretation

The authors propose and argue for a four-module, human-in-the-loop workflow hypothesis for embedding LLMs in TRA: data search, information extraction, integrated analysis and summarization, and expert interpretation, with each module as an expert-reviewed checkpoint and PoD selection and modifying-factor application retained by toxicologists. Relative to treating LLMs as end-to-end question answerers or as pure extraction tools, the work translates case observations into a process structure aligned with the draft ICH Q3E guideline and with regulatory expectations for transparency and traceability, and explicitly separates data retrieval from extraction and summarization. Based on an exemplar case study of oral PDE derivation for ATBC (GPT-5/GPT-5.2 accessed via ChatGPT, including Auto and Deep Research modes, repeated across independent runs, different users, and incognito sessions) interpreted against the available literature; the authors state it is not a statistical benchmarking study, quantitative performance comparison, or formal validation exercise.

The case indicates that the usable value of LLMs in TRA centers on pattern recognition, targeted retrieval, and rapid summarization: they accelerated extraction from full-text PDFs, produced structured tables (study identifiers, exposure route and duration, species and sex, endpoints, NOAEL/LOAEL values, and direct citations), and helped surface potentially overlooked studies. The work extends existing discussion of LLM extraction and evidence synthesis to how extracted evidence is combined and interpreted to derive a compound-specific limit, and it yields auditable tabular intermediates. Observational results from the modular prompting in the case study, with the authors noting that rigorous comparisons with human extractors remain limited.

End-to-end prompting was unstable: repeated runs with the same prompt changed the cited evidence, PoD selection, and modifying factors, producing oral PDE values of about 5, 20, 100, and 300 mg/day; even when the model was constrained to uploaded documents, PoD still varied between 100 and 300 mg/kg bw/day and PDE values varied at about 14, 25, and 100 mg/day. The work attributes this variability both to model stochasticity and to the expert-judgment-dependent non-uniqueness of TRA, and notes that document grounding alone is insufficient because the same evidence base can support multiple conclusions. Qualitative observations across multiple independent runs (including different users and incognito sessions); no systematic comparison across prompt configurations, models, or deployment modes was performed, and hallucination occurrence was not quantified.

The authors provide a mitigation table for the identified limitations: a structured retrieval protocol anchoring responses to the same verified evidence set, document-by-document processing before summarization, mandatory citation of each claim to primary documents, and logging of tool, model, version, date/time, and knowledge cutoff to support traceability and audit. These measures turn human-in-the-loop from a principle into operational control points, and stress that in data-rich settings retrieval must be separated from extraction and summarization, while in data-poor settings they may be combined into a single end-to-end prompt that still requires dedicated retrieval verification. Advisory proposals derived from case observations and literature synthesis; the authors note the modular workflow currently still requires expert review, source verification, and stepwise control, so efficiency gains may be partly offset by review and correction effort.

Perspective

The work is aimed at toxicological risk assessment practitioners, regulatory science researchers, and teams deriving compound-specific limits such as PDEs; it applies to data-rich substances assessed under the draft ICH Q3E framework and to similar evidence-intensive evaluations. It offers a workflow hypothesis and a mitigation checklist for confining LLMs to well-defined subtasks such as retrieval, extraction, and summarization, while leaving PoD selection, modifying-factor application, and final conclusions to human toxicologists. The authors note that in data-poor settings search and extraction may be combined into a single end-to-end prompt but still require dedicated retrieval verification, and that the next step is applying the workflow to more data-rich and data-poor substances and re-evaluations, with early engagement of regulatory bodies.

Readers should note that this is an incomplete reading: Figures 1 to 3, the prompt examples, extracted-evidence examples, and benchmark snapshots in the Supplementary Material are not reproduced in the main text, so exact prompt wording and run-by-run details cannot be checked; hallucinations are described as isolated and were not quantified; the case covers one data-rich substance (ATBC) and mainly GPT-5/GPT-5.2, with Claude Sonnet 4.5 and Gemini 3 used only as exploratory comparisons; and the authors note that efficiency gains may be partly offset by review and correction effort, leaving the net benefit and long-term reliability of the modular workflow in real regulatory processes an open question.

Sources