AI-RADS let 5 radiologists grade 350 AI outputs, reaching interreader agreement of α=0.87 for image tasks and α=0.93 for generative tasks
Synopsis
The study developed and multireader-evaluated AI-RADS, a structured framework for case-level assessment of radiology AI output reliability, clinical utility, and recommended actions, in which 5 board-certified radiologists independently evaluated 350 cases processed by 7 representative AI applications, assigning each case one of 5 AI-RADS categories, applicable modifiers, and an independent correctness rating as a reference; substantial interreader agreement was observed for core categories in image-based tasks (Krippendorff's α=0.87; 95% CI: 0.83-0.91) and generative AI tasks (α=0.93; 95% CI: 0.91-0.
Interpretation
The work proposes AI-RADS, a structured framework for case-level assessment of AI output reliability, clinical utility, and consequences for report communication, comprising 5 core categories plus applicable modifiers. The background states that despite the growing number of AI-based applications in radiology, no structured framework previously existed to assess their case-level reliability or to document overridden outputs. The framework was tested in a retrospective multireader study in which 5 board-certified radiologists independently rated 350 cases, each also receiving an independent correctness rating as a reference.
Core AI-RADS categories achieved substantial interreader agreement in both image-based and generative AI tasks. This provides a quantitative basis for the framework's reproducibility rather than a conceptual proposal alone. Krippendorff's α=0.87 (95% CI: 0.83-0.91) for image-based tasks and α=0.93 (95% CI: 0.91-0.95) for generative AI tasks, based on 5 readers and 350 cases.
Reader-assigned correctness corresponded directionally with AI-RADS categories: correct outputs mapped to categories 1 to 2, suitable for integration into clinical workflows, while outputs rated incorrect predominantly fell into categories 4 to 5, warranting override or removal from display. This links the category scale to an independent correctness reference, indicating the grading is not only consistent but also related to whether output is usable. Each case's independent correctness rating served as the reference against which reader-assigned AI-RADS categories were compared.
AI-RADS demonstrated applicability across 7 representative AI applications spanning image-based and generative tasks. The evaluation covers multiple task types rather than a single application or a single task. 350 cases from 7 representative AI applications were evaluated by 5 board-certified radiologists.
Perspective
The framework targets case-level assessment of AI output in radiology, applies to image-based and generative AI tasks, and is designed so that readers reach consistent judgments about output reliability, clinical utility, and recommended actions while overridden outputs can be documented in reports. The text shows it is operable across 7 representative AI applications and 350 cases, so next steps include routine institutional documentation of AI output, reader training, and cross-application comparison; beneficiaries include reporting radiologists, teams responsible for AI governance and quality control, and clinicians who receive AI information in reports.
The loaded text is an abstract-level full text: it does not provide the specific definitions of the 5 categories, the modifier list, the names and task details of the applications, the case distribution across categories, or how the correctness reference was determined, so readers who want to apply the framework directly still need the complete methods. In addition, interreader agreement was measured with 5 board-certified radiologists on a single retrospective dataset and 7 applications; agreement across other reader groups, institutions, and newer generations of AI applications remains an open question, and the alignment of categories 1 to 2 with correctness and of categories 4 to 5 with incorrect ratings is described directionally, without per-category numeric values in the text.
