Prospective Shadow-Mode Evaluation of an Artificial Intelligence Tool for Intracranial Aneurysm Detection on CT Angiography: Incremental Yield and Operational Impact
Synopsis
This prospective shadow-mode study evaluated an FDA-cleared AI algorithm (Aidoc) for intracranial aneurysm detection on 3,856 consecutive brain CT angiographies (November 7 to December 19, 2023) with radiologists blinded to AI, finding that AI alone achieved higher sensitivity (0.846) than radiologists alone (0.718) with similar specificity (0.987 vs 0.985), radiologist-AI concordance was 96.3%, AI surfaced additional aneurysms missed by radiologists with a relative enhanced detection rate (rEDR) of 39% (55 AI-only true-positives per 140 radiologist true-positives), an AI:radiologist incremental detection ratio of 1.83 (55 of 30), a favorable gain-to-pain ratio (GPR) of 1.20 (55 of 46), and a number-needed-to-examine (NNE) of 70.
Interpretation
In a real clinical workflow, AI alone detected intracranial aneurysms with higher sensitivity than radiologists alone, while specificity was similar between the two. Prior evaluations of AI aneurysm detection often relied on retrospective data or non-consecutive samples; this study used a prospective shadow-mode design on consecutive brain CTAs, ran the algorithm while radiologists were blinded to AI, and extracted report results via natural language processing, yielding sensitivity and specificity for both AI and radiologists within the real workflow. Based on 3,856 consecutive CTAs, AI sensitivity was 0.846 (0.787-0.894) versus radiologist 0.718 (0.649-0.779), with specificity of 0.987 (0.983-0.991) and 0.985 (0.981-0.989) respectively; non-overlapping confidence intervals indicate a stable sensitivity difference, and discordances were adjudicated by neuroradiologists with AI unblinding, with metrics estimated under a hybrid reference standard in which only discordances are adjudicated.
AI surfaced additional aneurysms beyond radiologist interpretation, suggesting that combining AI with radiologists can raise overall sensitivity. The study not only reported concordance but quantified incremental yield: the relative enhanced detection rate (rEDR) was 39%, i.e., 55 AI-only true-positives per 140 radiologist true-positives, and the AI:radiologist incremental detection ratio was 1.83 (55 of 30), favoring AI. The increment rests on a 96.3% radiologist-AI concordance (3,714 of 3,856) plus adjudication of discordances; most AI-only aneurysms were under 3 mm (31 of 55, 56.4%) and radiologist-only ones were mostly 3 to 5 mm (18 of 33, 54.5%), indicating the increment concentrates in smaller lesions.
Operationally, the gain-to-pain ratio and number-needed-to-examine were generally acceptable, but varied markedly across care settings. The study introduced operational metrics such as GPR and NNE, placing detection benefit and additional workload in one framework and reporting them by setting, which is less common in prior AI evaluations focused mainly on diagnostic accuracy. Overall GPR was 1.20 (55 of 46) and NNE was 70.1 (3,856 of 55); metrics were most favorable in inpatients (rEDR 78.3%, GPR 2.57, NNE 28.9), intermediate in emergency (rEDR 37.1%, GPR 1.0, NNE 85.3), and unfavorable in outpatients (rEDR 14.1%, GPR 0.67, NNE 130.3).
Perspective
These results apply to consecutive brain CTAs in a shadow-mode setting where radiologists were blinded to AI, report results were extracted via natural language processing, and only discordances were adjudicated; they primarily serve imaging and clinical teams assessing AI-assisted aneurysm detection, and the more favorable operational metrics in inpatient and emergency settings point to setting-specific deployment and workflow design as the most direct application of this evidence.
A careful reader would still watch: how the hybrid reference standard with adjudication of only discordances affects sensitivity and specificity estimates; what the clinical significance and follow-up value are of AI-only lesions that are mostly under 3 mm; how the workload burden reflected by outpatient GPR 0.67 and NNE 130.3 evolves over longer operation; and how reproducible the findings are across other institutions, other AI tools, and different populations given a single institution, a single algorithm, and a specific time window. This reading is at the abstract level and does not include figures or supplementary material, so the original should be consulted to verify stratified details and statistical methods.
