Evaluating Milvue SmartUrgences for Hip Fracture Diagnosis on Emergency Department Radiographs
Synopsis
This retrospective diagnostic accuracy study enrolled 539 consecutive patients aged 60 years or older undergoing hip and pelvis radiography at two emergency departments, used a consensus panel of three musculoskeletal radiologists and two senior orthopaedic trauma surgeons blinded to AI output as the reference standard, and evaluated Milvue SmartUrgences, which classified radiographs as "YES," "NO," or "DOUBT"; 487 patients received definitive classifications and entered the primary analysis (mean age 83.4 years, SD 8.4; 66% female; 191 hip fractures identified by the reference standard), yielding sensitivity 97.4% (95% CI 94.0-98.9), specificity 86.5% (95% CI 82.1-89.9), positive predictive value 82.3% (95% CI 76.8-86.7), negative predictive value 98.1% (95% CI 95.6-99.2), accuracy 90.
Interpretation
In a consecutive emergency department cohort, Milvue SmartUrgences showed high sensitivity (97.4%) and high negative predictive value (98.1%) for hip fracture detection. Prior evaluations of AI hip fracture detection have often drawn on non-consecutive or single-centre data; this study used consecutive patients aged 60 years or older at two emergency departments and reported a full set of sensitivity, specificity, predictive values, and accuracy. Retrospective diagnostic accuracy study with 539 patients enrolled and 487 in the primary analysis; the reference standard was a consensus panel of three musculoskeletal radiologists and two senior orthopaedic trauma surgeons blinded to AI output, and each metric is reported with a 95% confidence interval.
The software produced a three-way output of "YES," "NO," or "DOUBT," with the primary analysis restricted to definitive classifications and "DOUBT" evaluated separately. This design separates uncertain outputs from the main performance estimates, letting readers distinguish definitive readings from cases requiring human review. Of 539 patients, 487 received definitive classifications, indicating that a subset was labelled "DOUBT" and handled separately; the text does not report the number or performance for that subset.
The positive predictive value of 82.3% was lower than the negative predictive value of 98.1%, indicating a remaining gap between a positive AI flag and confirmed diagnosis in this cohort. The study reports PPV and NPV alongside sensitivity and specificity, allowing readers to interpret AI alerts against the actual disease prevalence in an emergency setting. PPV 82.3% (95% CI 76.8-86.7) and NPV 98.1% (95% CI 95.6-99.2) were calculated against 191 hip fractures confirmed by the reference standard.
The authors conclude that AI output should complement rather than replace clinical and radiological assessment. This framing positions the study as an evaluation of a decision-support tool rather than a validation of a replacement for radiologist interpretation. The conclusion follows directly from the cohort's sensitivity, specificity, predictive values, and accuracy and represents the authors' interpretation of those data.
Perspective
The results address patients aged 60 years or older undergoing hip and pelvis radiography in the emergency department, and the main performance estimates rest on the 487 patients who received a definitive "YES" or "NO" classification; the study does not directly provide performance conclusions for "DOUBT" cases, other age groups, non-emergency settings, or other imaging equipment and institutions. As a retrospective design, data from two emergency departments can serve as a reference for subsequent prospective evaluation and workflow integration studies.
Readers may still wonder about the number and performance of "DOUBT" classifications, how the software performs across different institutions and equipment, and how far a retrospective design generalizes to prospective emergency workflows; because this text is presented as a summary without figures or subgroup detail, these questions remain open in the available material.
