U-Net-Based Automated Quality Control of Knee Radiographs: Dual-Center Validation and Clinical Intervention
Synopsis
This study developed and dual-center validated an interpretable U-Net-based AI framework for automated quality control (QC) of knee anteroposterior (AP) and lateral (LAT) radiographs, generating QC indices through anatomical segmentation and landmark localization, with mean Dice similarity coefficients of 0.964 and 0.936 in internal and external validation, most intraclass correlation coefficients exceeding 0.90, QC sensitivity of 91.67% to 98.36% and specificity of 89.13% to 99.10%; separately, six radiographers who received 4 weeks of AI-based feedback on 948 radiographs from 474 patients showed exploratory numerical gains in sensitivity, with effect sizes of 0.30 to 0.73.
Interpretation
The study built an interpretable AI QC framework that converts anatomical segmentation and landmark localization into QC indices for knee AP and LAT radiographs. Relative to current QC that mainly relies on subjective, inefficient manual assessment, this work automates the QC workflow and provides interpretable quantitative indices. Developed and validated on 1600 adult single-knee AP and LAT radiographs from 800 patients at two centers; segmentation was evaluated by Dice similarity coefficient, with mean DSCs of 0.964 and 0.936 in internal and external validation.
The framework showed high QC discrimination in dual-center validation, with most agreement indices exceeding 0.90. Existing automated tools are described as lacking sufficient clinical validation, whereas this study provides validation results in both internal and external cohorts. QC sensitivity ranged from 91.67% to 98.36% and specificity from 89.13% to 99.10%, with most ICCs exceeding 0.90; landmark localization was assessed by Euclidean error, normalized distance error, and percentage of correct keypoints, with PCK@20% of 100.0%/99.4% for AP and 81.6%/61.8% for LAT (internal/external).
The study explored the clinical value of AI feedback on radiographers' QC performance, with exploratory numerical gains in sensitivity after feedback. Beyond model validation, it examined the potential impact of integrating AI QC into the actual workflow. Six radiographers received 4 weeks of AI-based feedback covering 948 radiographs from 474 patients, with effect sizes of 0.30 to 0.73; the authors frame this as an exploratory trend and note that prospective controlled studies are needed.
Perspective
The framework targets automated QC of adult single-knee knee AP and LAT radiographs, suited to radiology workflows that have dual-center validation conditions and wish to replace or supplement manual QC with objective indices; the workflow evaluation involved 6 radiographers and 4 weeks of AI feedback on 948 radiographs from 474 patients, an exploratory setting whose conclusions require confirmation in prospective controlled studies.
In external validation, LAT landmark localization error was higher than AP and PCK@20% dropped noticeably, suggesting that performance differences across centers and projection views deserve attention; the post-feedback sensitivity gains are exploratory numerical results with effect sizes of 0.30 to 0.73 and require prospective controlled studies to confirm robustness; moreover, this is abstract-level information, so implementation details of segmentation and localization, threshold settings for QC indices, and the operational design of the feedback mechanism still require consulting the original figures and full text.
