PC-Seg lifts sparse 2D annotations to 3D OCT segmentation via five-stage curriculum learning, matching full supervision with about 0.7% of labels
Synopsis
The work proposes PC-Seg (Progressive Cross-view Segmentation), a five-stage curriculum learning framework in which a single 2D model first learns cross-view consistency between standard B-scans and orthogonal slices to generate reliable volumetric pseudo-labels, which are then distilled into a 3D model and followed by 2D/3D co-training with ensemble pseudo-labeling; on the public MSHC and Duke DME OCT datasets it reaches segmentation accuracy comparable to fully supervised learning using only about 0.7% of the labeled data, outperforming the semi-supervised and retinal layer segmentation methods it compares against.
Fig. 1. Overview of the five-stage curriculum learning in PC-Seg. The 2D model trained in Stage 3 is utilized again in Stage 5 for the final ensemble. The “Train” arrows in each stage denote the optimization using the baseline SSL framework in Sect. 2.1.
· Page 3Interpretation
It introduces a five-stage curriculum that progressively lifts sparse 2D annotations to a 3D segmentation model: Stage 1 warms up a 2D model on B-scans, Stage 2 adds orthogonal B-scans, Stage 3 performs cross-teaching between orthogonal planes, Stage 4 distills 3D pseudo-labels into a 3D model, and Stage 5 runs 2D/3D co-training. Unlike multi-view approaches that rely on annotations on orthogonal planes or dense voxel-wise labels, the method uses a single 2D model to learn cross-view consistency, accommodating OCT anisotropy and sparse annotation conditions. The paper presents the five-stage pipeline in Fig. 1 and reports an ablation on MSHC with 6 labeled images, where the F-score rises from 0.8744 at Stage 1 to 0.9047 for the Stage 5 ensemble prediction.
Cross-view consistency and orthogonal slices as additional unlabeled data yield concrete gains: introducing orthogonal B-scans in Stage 2 raises the B-scan F-score from 0.8744 to 0.8917, and the Stage 3 cross-teaching further narrows the performance gap between the two planes. In the ablation, the 'Only stage 3' variant that starts cross-teaching without sufficient prior B-scan learning reaches F-scores of 0.8713/0.8748, below the full pipeline's Stage 3 (0.8952/0.8969), indicating that progressively expanding dimensions and views is needed to curb pseudo-label error accumulation. Evidence comes from the MSHC ablation table with 6 labeled images, including the 'Only stage 3' and 'w/o ortho B-scan' controls that skip progressive steps.
On MSHC, PC-Seg outperforms the compared 2D and 3D semi-supervised baselines across all three labeled settings (6, 30, and 60 images); with 60 labeled images (about 0.7% of training data) the ensemble reaches an F-score of 0.9174, close to fully supervised methods trained on all 8,820 images (e.g., Structured-Layer 0.9193, 1D+2D U-Net 0.9198). The standard 3D semi-supervised baseline (3D ResUNet w/ UM) improves only modestly under sparse labels (0.7843 with 6 labeled images), whereas the progressive curriculum bridges 2D and 3D learning. Evidence is the quantitative comparison table on MSHC covering supervised methods, domain-specific semi-supervised methods, and general semi-supervised frameworks.
On the Duke DME dataset, the method markedly improves the Fluid class, reaching an F-score of 0.742 in the ensemble prediction, while the compared conventional methods sit around 0.3–0.6 on Fluid; the qualitative comparison shows the 3D model using spatial context to smooth striping artifacts of orthogonal 2D predictions. This dataset has scarce labels and highly diverse fluid lesions, and incorporating 3D spatial context improves the completeness of complex pathological structure detection. Evidence is the quantitative table on Duke DME trained with 66 labeled B-scans (11 slices × 6 subjects) plus the qualitative comparison in Fig. 2.
Perspective
The result targets retinal OCT volumetric layer and lesion segmentation, in settings where only a few standard B-scans are labeled and orthogonal slices are available as unlabeled data; the framework is described as model-agnostic, so it could in principle be combined with other 2D/3D backbones, but the experiments here are limited to ResUNet (depth=4, base_channels=32) and the two public datasets MSHC and Duke DME. For ophthalmic imaging groups aiming to cut annotation cost, the pipeline offers a reusable curriculum design.
The ablation is run only under the 6-labeled-image setting on MSHC, so whether each stage contributes consistently at other label budgets is unclear; lowering the lesion-class confidence threshold to 0.5 in Stages 4 and 5 relies on the assumption that ensemble predictions are more stable, and its applicability across lesion types and datasets remains to be observed; the paper also does not report inference time or memory cost, so deployment cost needs separate evaluation.
