Skip to main content
Back to timeline
medRxivSource publication:

Task-Specific Quality Gating for Retinal OCT B-Scans: Learned Representations Over Scalar Metrics in Choroid Segmentation

Synopsis

Using choroid segmentation as a prototype task on 6,076 OCT B-scans from 80 subjects, this study systematically compared scalar no-reference image quality metrics (BRISQUE, NIQE, PIQE, SNR, PSNR), general-purpose ImageNet-pretrained representations, and retinal foundation models as task-specific quality gates, finding that scalar metrics correlate weakly with segmentation Dice (|r| < 0.20), that general-purpose pretrained representations reach linear-probe ROC-AUC up to about 0.77, that the OCT-specific foundation model RETFound reaches about 0.81, and that only the retinal foundation model embeddings form quality-aligned unsupervised K-Means clusters exceeding a patient-level permutation null.

AI-generated editorial illustration: Task-Specific Quality Gating for Retinal Optical Coherence Tomography B-Scans: Learned Representations Over Scalar Metrics in Choroid Segmentation

Interpretation

Scalar no-reference image quality metrics fail to predict downstream segmentation reliability: BRISQUE r = -0.141, NIQE r = -0.030, PIQE r = -0.081, SNR r = -0.194, PSNR r = -0.072, all with |r| < 0.20. Device-reported signal-strength indices were already known to be weak predictors, but the ability of natural-scene-statistics no-reference metrics such as BRISQUE, NIQE, and PIQE to predict downstream OCT segmentation reliability had not been directly tested; this work provides that direct evaluation. Pearson correlations computed on the full 6,076-image dataset, a large sample, though observational and limited to a single device and a single segmentation model.

The failure of NR-IQA lies largely in the aggregation step: the 36-dimensional BRISQUE and NIQE feature vectors reach linear-probe ROC-AUC of 0.685 and 0.674, above chance. This indicates that these hand-crafted features themselves carry information relevant to segmentation quality that is lost when compressed into a single scalar score, offering a new perspective on how such features might be used for quality gating. Leave-one-patient-out cross-validation across 80 subjects with patient-level bootstrap confidence intervals; BRISQUE 0.685 [0.618, 0.743] and NIQE 0.674 [0.607, 0.735].

Learned representations outperform scalar metrics, and domain-specific pretraining adds further benefit: ImageNet-pretrained backbones reach ROC-AUC up to 0.773 (ViT-B/16), while retinal foundation models reach 0.809 (RETFound) and 0.803 (UrFound), with the foundation-model group exceeding the ImageNet group overall (delta ROC-AUC = 0.051, 95% CI [0.023, 0.080], p < 0.001). This is a systematic comparison, on a single benchmark, of task-specific quality gating across three tiers (scalar metrics, general-purpose pretrained representations, retinal foundation models), and it suggests that domain relevance matters more than the specific pretraining supervision recipe, since UrFound and RETFound do not differ significantly (delta ROC-AUC = 0.006, p = 0.43). Pooled out-of-fold predictions from 80 LOPO folds, paired patient-level bootstrap tests (5,000 resamples), plus randomly initialized backbone controls to separate architecture from pretrained weights.

In unsupervised geometry, only the retinal foundation model embeddings align with quality: UrFound (k* = 3) and RETFound (k* = 3) show Bad-rate spreads across clusters of 48.9 and 47.1, exceeding patient-level permutation nulls of 15.8 and 15.6 (p = .009 and p = .010), whereas ResNet-50 and EfficientNet-B0 do not exceed their nulls and ViT-B/16 falls short of significance (p = .073). This shows that domain-specific pretraining not only improves linear separability but also makes segmentation quality a geometrically salient unsupervised axis of the representation, while generic ImageNet pretraining at fixed k = 2 significantly reduces quality-aligned separation relative to random initialization (ResNet-50 delta Spread = -19.7, ViT-B/16 -24.9, both p < 10^-30). 300 patient-level resamples, Kneedle selection of k*, patient-level permutation null tests, and reported 95% intervals.

Perspective

The results apply to a setting with a single OCT device, a healthy subject population, and a single Residual U-Net choroid segmentation model, with quality labels operationalized by a Dice threshold of 0.85 that the authors explicitly note is not an intrinsic boundary of image quality and that the framework can be evaluated under alternative thresholds. The direct implication is that a reliable choroid quality gate does not require bespoke architecture design: a frozen retinal foundation model backbone with a supervised linear head trained on Dice already achieves ROC-AUC of 0.809, and the critical design requirements are that quality labels be task-specific (derived from segmentation model performance rather than human readability) and that the backbone be pretrained on the target modality. This paradigm provides a basis for evaluating whether learned representations can support pre-inference quality assessment in other medical imaging tasks.

The authors note that a causal signal has not been established and that the results are encouraging rather than causal; the study evaluates pretraining data rather than the training objective, a distinction left for future work; image attributes such as choroidal boundary visibility are not explicitly evaluated; and few-shot quality gating with limited task-specific data remains to be explored. In addition, the loaded text is a fast parse in which the full numerical detail of Figures 3 and 4 is not given point by point, so the threshold-sweep specifics can only be understood from the prose description, which may limit fine-grained appreciation of the threshold-sensitivity conclusions.

Sources