Skip to main content
Back to timeline
PloS oneSource publication:

Deterministic and stochastic interventions in reducing drug-drug interactions in inappropriate prescribing: A systematic review

Synopsis

This systematic review searched PubMed, Scopus, ScienceDirect, and IEEE Xplore following PRISMA 2020 and the SPIDER framework, included 10 primary studies of computational and clinical decision support interventions aimed at reducing drug-drug interactions or inappropriate prescribing, synthesized them narratively across deterministic, ontological, and stochastic/generative categories, and assessed risk of bias with PROBAST+AI, finding that earlier deterministic systems showed modest improvements in prescribing process measures with inconsistent links to patient-level outcomes, that recent stochastic and generative models reported strong internal performance metrics, and that AI-driven studies carried a consistently high risk of bias in the analysis domain driven mainly by limited external

Source-provided article image: Deterministic and stochastic interventions in reducing drug-drug interactions in inappropriate prescribing: A systematic review.
PubMed

Interpretation

The review divides DDI mitigation research into deterministic (including ontological) and stochastic/generative mechanisms, classifying studies by the underlying decision-making mechanism reported by their authors rather than by clinical domain or stated aim. Prior syntheses tended to examine health-system interventions or computational prediction models in isolation; this review brings policy-level, clinical operational, and emerging AI-driven approaches into a single analytical frame. Based on narrative synthesis of 10 included studies, with searches across four databases from inception to March 2026 and screening and extraction performed independently by two reviewers with third-reviewer arbitration.

PROBAST+AI assessment showed a mixed rather than predominantly favorable risk-of-bias profile: three studies low risk, three unclear, and four high risk, with high risk concentrated among studies using stochastic or generative architectures and most commonly driven by the Analysis domain. The review incorporated AI- and machine-learning-specific sources of bias, including overfitting, inappropriate handling of missing data, data leakage between training and validation sets, and adequacy of model validation, and applied the worst-domain-anywhere principle for overall judgment. Assessed across the four PROBAST+AI domains of Participants and Data Source, Predictors, Outcome, and Analysis, with study-level signaling-question evaluations provided in Supporting Information S2 Table.

The review identifies highly heterogeneous performance metrics as a cross-cutting finding: institutional workflow systems report alert acceptance and prescribing error counts, semantic frameworks report ontology coverage and class-mapping accuracy, and data-driven predictive models report AUC, F1-score, Jaccard similarity, or RMSE. The review treats metric heterogeneity itself as the central story of the field's evolution rather than merely a methodological inconvenience, and notes that no single evaluation framework has kept pace across the full span of technical evolution. Because outcome metrics are disjoint, no quantitative meta-analysis was performed and no sensitivity analysis was conducted, so cross-study comparison is qualitative and interpretive.

The review points to a disconnect between computational prediction performance and measurable clinical endpoints: early deterministic workflow interventions framed utility in terms such as preventable adverse drug events and QALYs saved, whereas recent stochastic and generative models largely report internal metrics without extending evaluation to real-world clinical environments to measure corresponding reductions in adverse drug events or patient morbidity. The review frames this disconnect as the core of the 'last mile' problem in medical informatics and calls for a shared, translationally-oriented outcome framework capable of expressing both computational accuracy and clinical impact. Based on comparison of how included studies measured outcomes, with the authors stating explicitly that the current literature does not provide sufficient evidence to conclude that improved DDI prediction performance directly leads to better patient outcomes.

Perspective

The review's conclusions apply to researchers evaluating computational and clinical decision support interventions for reducing drug-drug interactions or inappropriate prescribing, to clinical decision support designers, and to health systems considering adopting such systems; its analytical frame covers deterministic, ontological, and stochastic/generative mechanisms, and included studies had to report an evaluation component such as improvements in prescribing accuracy, reductions in prescribing errors, performance metrics, or model validation. The review explicitly excluded tools used exclusively for diagnostic purposes, dispensing or administration errors, publications without empirical data, grey literature, and studies whose full text was unobtainable, so its conclusions are scoped to DDI mitigation at the prescribing stage. The authors note that a high PROBAST+AI rating reflects a model's stage of clinical translational readiness at the time of publication rather than an inherent or permanent flaw in the underlying model, and they expect that rating to improve once the Participants and Data Source domain is addressed through representative, adequately documented patient populations and outcome applicability is addressed through prospective clinical validation against real-world patient outcomes.

A careful reader would still watch several open questions: because only 10 studies were included and their outcome metrics are disjoint, cross-generational performance comparison can only be qualitative, and whether conclusions would change if a shared outcome framework expressing both computational accuracy and clinical impact emerged remains to be seen; external validation, calibration reporting, and handling of overfitting and data leakage remain limited in the included stochastic and generative studies, and future reporting on these points will shape judgments of their clinical readiness; whether the geographic and care-standard bias introduced by reliance on datasets such as MIMIC-III is mitigated in multi-region, multi-population cohorts still needs verification; and although this was a full-text reading, the study-characteristics table, the risk-of-bias traffic-light plot, and the item-level PROBAST+AI assessment reside in figures and supporting information, so readers checking domain-level judgments for individual studies should consult Table 1, Fig 2, and S2 Table in the original.

Sources