Oracle-curated synthetic fundus images lift ConvNeXt diabetic retinopathy grading kappa from 0.6350 to 0.7695 and double proliferative DR recall
Synopsis
The study pretrained StyleGAN3 and Medfusion on an auxiliary dataset, curated synthetic images through a domain-matched Oracle network using Latent Space Rejection Sampling and Class-Conditioned SDEdit Escalation, and evaluated downstream grading on the Indian DR Image Dataset across ConvNeXt, ResNet50, and VGG16 with five random seeds, finding that a real-data-only ConvNeXt baseline reached a quadratic weighted kappa of 0.6350 ± 0.0135, curated StyleGAN3 augmentation raised it to 0.7185 ± 0.0309, Medfusion augmentation reached 0.7695 ± 0.0254 (an absolute improvement of 0.1345), and Medfusion doubled proliferative DR recall from 0.2615 ± 0.1595 to 0.5385 ± 0.0942.
Interpretation
Strictly curated synthetic images improve DR grading under limited data, with Medfusion augmentation raising ConvNeXt quadratic weighted kappa from 0.6350 ± 0.0135 to 0.7695 ± 0.0254, an absolute gain of 0.1345. Prior synthetic augmentation for DR grading was often limited by mode collapse and severity downshift; this work introduces a domain-matched Oracle network with Latent Space Rejection Sampling and Class-Conditioned SDEdit Escalation to curate images, directly linking generation quality to downstream ordinal classification performance. Evaluated on an isolated downstream DR classification task (Indian DR Image Dataset) across three architectures (ConvNeXt, ResNet50, VGG16) using five random seeds, with means and standard deviations reported.
Medfusion augmentation doubled proliferative DR recall from 0.2615 ± 0.1595 to 0.5385 ± 0.0942, i.e., minority-class sensitivity. Most synthetic-data work focuses on overall accuracy; this work explicitly reports sensitivity for the most severe and least represented proliferative DR class, directly addressing the clinical pain point of data imbalance. Reported as means and standard deviations over five random seeds and compared against a real-data-only baseline on the same isolated task.
Global image quality metrics such as FID may not reflect localized pathological utility: curated StyleGAN3 achieved better FID (24.25) and padded structural similarity (0.9037), whereas Medfusion had FID 58.70 and padded structural similarity 0.8852 but a higher inception score (2.35), yet Medfusion ultimately delivered superior downstream clinical performance. This suggests that ranking synthetic medical images by FID alone is insufficient and should be combined with downstream task performance and structural diversity. Reports three generation metrics (FID, padded structural similarity, inception score) alongside downstream classification metrics, enabling direct comparison.
Perspective
The result targets automated DR grading under data scarcity and class imbalance, suited to research or engineering teams that have an auxiliary dataset for pretraining generative models and can build a domain-matched Oracle for curation; its value lies in providing additional training signal for the minority proliferative DR class, supporting faster and more sensitive automated diagnostic workflows.
The curation pipeline depends on a domain-matched Oracle network and an auxiliary dataset, so its construction cost and reproducibility on other datasets remain to be seen; whether the FID-versus-downstream-performance divergence holds across more diseases and imaging modalities, and whether the proliferative DR recall gain persists on larger, multi-center data, are open questions worth further validation.
