Public articles linked to the same research event.
Journal of the American Medical Informatics Association : JAMIA This study developed and validated DiagnosticXchange, an open-source clinical simulation framework in which AI systems diagnose cases by ordering tests, requesting imaging, and performing procedures, with each action mapped to CPT codes capturing cost, time, work relative value units, and invasiveness; using 8 large language models on 216 peer-reviewed cases across 19 specialties (1728 sessions), it found that three systems with near-identical accuracy (93.5%–94.0%) differed significantly in cost (P < .001; 1.75-fold between the most and least expensive) and 2.
This study developed and validated DiagnosticXchange, an open-source clinical simulation framework in which AI systems diagnose cases by ordering tests, requesting imaging, and performing procedures, with each action mapped to CPT codes capturing cost, time, work relative value units, and invasiveness; using 8 large language models on 216 peer-reviewed cases across 19 specialties (1728 sessions), it found that three systems with near-identical accuracy (93.5%–94.0%) differed significantly in cost (P < .001; 1.75-fold between the most and least expensive) and 2.
This study developed and validated DiagnosticXchange, an open-source clinical simulation framework in which AI systems diagnose cases by ordering tests, requesting imaging, and performing procedures, with each action mapped to CPT codes capturing cost, time, work relative value units, and invasiveness; using 8 large language models on 216 peer-reviewed cases across 19 specialties (1728 sessions), it found that three systems with near-identical accuracy (93.5%–94.0%) differed significantly in cost (P < .001; 1.75-fold between the most and least expensive) and 2.
This study developed and validated DiagnosticXchange, an open-source clinical simulation framework in which AI systems diagnose cases by ordering tests, requesting imaging, and performing procedures, with each action mapped to CPT codes capturing cost, time, work relative value units, and invasiveness; using 8 large language models on 216 peer-reviewed cases across 19 specialties (1728 sessions), it found that three systems with near-identical accuracy (93.5%–94.0%) differed significantly in cost (P < .001; 1.75-fold between the most and least expensive) and 2.