Retrospective comparison of three commercial artificial intelligence algorithms for detection of intracranial hemorrhage (ICH) in the emergency radiology department
Synopsis
This retrospective study analyzed 4027 consecutive non-contrast head CT examinations from a large emergency hospital in southwest Sweden to compare three commercial AI algorithms for ICH detection, finding substantial variation with only Aidoc demonstrating clinically relevant accuracy (90.3% sensitivity, 99.0% specificity), while a simulated mathematical combination of Aidoc with a human reader increased sensitivity to 96.0% while maintaining 99.4% specificity (P < 0.001), comparable to two radiologists.
Interpretation
Of 3902 evaluable examinations, 176 cases (4.5% prevalence) were confirmed as ICH by two-tier consensus adjudication and 209 were excluded, with eight ICH cases missed by both radiologists detected by at least one AI system. The study independently compared three commercial algorithms on the same consecutive emergency NCHCT dataset, using expert manual review with two-tier consensus adjudication as the reference standard rather than relying on vendor reports or a single reader. Based on a retrospective design of 4027 consecutive examinations, with all positive or discrepant cases undergoing two-tier consensus adjudication review, providing a relatively rigorous reference standard; however, it is single-center retrospective data.
Algorithm performance varied substantially, with only Aidoc demonstrating clinically relevant accuracy, at 90.3% sensitivity and 99.0% specificity. Independent clinical validation of commercial ICH detection algorithms has been limited; this study provides head-to-head comparison results across multiple algorithms on the same dataset. Results come from direct comparison on the same consecutive examinations and report specific metrics such as sensitivity and specificity; the study is retrospective and single-center.
A simulated mathematical combination of Aidoc with a human reader increased sensitivity to 96.0% while maintaining 99.4% specificity (P < 0.001), comparable to two radiologists. The study goes beyond evaluating a single algorithm by simulating human-AI combined detection performance, suggesting that joint interpretation can improve ICH detection. The combination assessment is based on an idealized logical OR model that assumes radiologists perfectly dismissed all false-positive AI flags to calculate system specificity, making it a simulation rather than real prospective workflow validation.
Perspective
This is a single-center retrospective study applicable to the setting of consecutive non-contrast head CT examinations at a large emergency hospital in southwest Sweden, aimed at emergency radiology readers interested in the relative performance of commercial ICH detection algorithms and the potential of human-AI collaboration; the human-AI combination result is based on an idealized logical OR model simulation rather than a real prospective workflow.
Readers may still wonder: whether performance would hold in real workflows, given that the idealized logical OR model assumes radiologists perfectly dismiss all false-positive AI flags; how single-center retrospective data generalize to other institutions; and what underlies the difference whereby only one of the three algorithms achieved clinically relevant accuracy.
