DiagnosticXchange: An Open-Source Framework for Evaluating Safety, Efficiency, and Diagnostic Reasoning in Clinical AI Systems
Synopsis
This study developed and validated DiagnosticXchange, an open-source clinical simulation framework in which AI systems diagnose cases by ordering tests, requesting imaging, and performing procedures, with each action mapped to CPT codes capturing cost, time, work relative value units, and invasiveness; using 8 large language models on 216 peer-reviewed cases across 19 specialties (1728 sessions), it found that three systems with near-identical accuracy (93.5%–94.0%) differed significantly in cost (P < .001; 1.75-fold between the most and least expensive) and 2.
Interpretation
It proposes and validates an open-source, reproducible multi-dimensional clinical AI evaluation framework that maps diagnostic actions to CPT codes to measure accuracy, cost, time, invasiveness, and physician effort simultaneously. Relative to existing accuracy-only benchmarks, the framework brings clinical resource consumption and safety behaviors into a single evaluation pipeline and releases it as an open-source tool. Validated with 8 large language models on 216 peer-reviewed cases across 19 specialties (1728 sessions), using competing-risk survival analysis, unsupervised clustering, diagnostic contribution scoring, and a pilot comparison with 14 neurologists.
Clinical AI systems with similar accuracy can differ in clinically meaningful ways in cost and physician oversight requirements. Three systems differed by only 0.5 percentage points in accuracy (93.5%–94.0%) yet showed a 1.75-fold cost difference (P < .001) and a 2.1-fold difference in physician oversight, showing that accuracy metrics cannot distinguish these differences. Based on paired sessions on the same case set, the cost difference was statistically significant (P < .001), and competing-risk analysis showed the most efficient system solved 86% of cases within $5000 while the most resource-intensive required $10000 for equivalent accuracy.
AI systems' reasoning behavior profiles predict resource consumption better than system identity and can be summarized by unsupervised clustering into a few strategies. Clustering identified 3 reasoning strategies, and behavioral profiles predicted resource consumption with an AIC of 4683 versus 5056 for system identity, suggesting evaluation should focus on behavioral patterns rather than model brand alone. Based on behavioral data from 1728 sessions with unsupervised clustering and model comparison, where the AIC difference supports better predictive power for behavioral features.
Safety analysis revealed clinically relevant risk behaviors invisible to accuracy metrics, including premature diagnosis, noncontributory invasive procedures, and futile invasive procedures on failed cases. It quantifies specific safety behaviors beyond accuracy, providing actionable dimensions for pre-deployment safety evaluation of clinical AI. Across sessions on 216 cases, it recorded premature diagnosis up to 9.3%, noncontributory invasive procedures 8.7%–29.9%, and futile invasive procedures on failed cases 3.7%–18.1%.
Perspective
The framework is intended for pre-deployment evaluation of clinical AI diagnostic systems, suitable for comparing different models or systems in a simulated hospital environment across accuracy, cost, time, invasiveness, and physician effort; its open-source nature allows other researchers to reuse it on their own case sets and models. The comparison with 14 neurologists is pilot in nature, suggesting the framework can support human-AI comparison, but the applicable scope of such comparison needs further clarification in larger populations and more scenarios.
Test-retest analysis showed reproducible accuracy (84.6% concordance) but substantial process variability (cost CV: 86.6%), suggesting that cost results from a single evaluation may be unstable and readers should attend to cost distributions across repeated runs. The comparison with 14 neurologists is a pilot, and the generalizability of its conclusions remains unclear. In addition, how differences between the simulated hospital environment and real clinical workflows affect cost, invasiveness, and safety behavior metrics remains an open question for further observation.
