Replacing human labels with seven models across four abdominal CT datasets removed pretraining's dependence on label quality while direct deployment stayed quality-sensitive
Synopsis
Across four abdominal CT datasets (WORD, AMOS, CT-1K, AbdomenAtlas), the authors generated pseudo-label variants with seven models (nnU-Net, MedSAM, TotalSegmentator, and four STU-Net sizes), then trained DynUNet under identical deterministic settings for in-domain training and for pretraining followed by fine-tuning; in-domain performance rose strongly and non-linearly with label quality and dataset volume partly compensated for quality, whereas fine-tuned models significantly outperformed a no-pretrain baseline in the vast majority of settings and pretraining label quality no longer clearly affected downstream results.
Interpretation
The study characterizes a pseudo-label quality spectrum: nnU-Net reaches the highest Dice (83.7% to 95.2%), MedSAM the lowest (Dice 39.4% and Surface Dice 24.8% on AbdomenAtlas), STU-Net improves from small to huge with diminishing gains, and TotalSegmentator sits in between. Prior label-noise work often used synthetic noise from morphological operations or perturbations, or focused on 2D and RGB modalities; this work uses realistic pseudo-label errors with native 3D models across four datasets and seven generators. Based on Dice and Surface Dice means and standard deviations in Table 1 and the ECDF curves in Figure 1; only anatomical structures predictable by all generators are averaged for fair comparison.
In-domain training shows a strong positive but non-linear quality-performance relationship, with diminishing returns at the high-quality end falling below the y=x reference; dataset volume appears to compensate for quality, as AbdomenAtlas (5k cases) plateaus early while smaller datasets such as Word and Amos remain sensitive to label improvements. It compares the quality-performance relationship across data scales, indicating that the marginal value of quality investment shrinks as dataset size grows. Scatter plots with second-degree polynomial regression in Figure 2, covering four datasets and both Dice and Surface Dice.
In the pretraining setting, label quality is no longer a primary factor: models pretrained on low-quality MedSAM labels and then fine-tuned still significantly outperform the no-pretrain baseline and reach scores comparable to pretraining on much higher-quality labels; under a fixed compute budget, pretraining dataset size (Word, Amos versus AbdomenAtlas) also shows no major impact. The authors describe this as the first benchmark isolating the relationship between pretraining label quality and downstream performance after fine-tuning, extending label-quality analysis from inference models to the pretraining stage. Global and organ-level Dice and Surface Dice on SegThor and FLARE in Figure 3, with significance from a one-sided Wilcoxon signed-rank test; the authors also extended baseline training from 1k to 2k steps to confirm gains are data-driven rather than from training duration.
The authors derive two actionable insights: for dataset creators targeting massive pretraining corpora, extensive expert refinement may not be cost-effective; for model developers, simple pseudo-labeling of unlabeled image collections can suffice to learn transferable representations. It shifts label-quality research from method fixes toward data-curation decisions, giving a basis for whether resources should go to pretraining corpora or downstream target datasets. The conclusion rests on consistent patterns across four datasets, seven generators, and two downstream target datasets, and is an empirical recommendation within this experimental setting.
Perspective
The study targets abdominal CT multi-organ segmentation with DynUNet and an nnU-Net-style training pipeline; pretraining labels come from nnU-Net, MedSAM, TotalSegmentator, and four STU-Net size variants, and downstream targets are SegThor and FLARE. Its conclusions apply to the workflow of large-scale pseudo-label pretraining followed by high-quality downstream fine-tuning, and to the contrasting case of direct deployment without manual annotations. For dataset creators, it suggests re-evaluating expert refinement investment in pretraining corpora; for model developers, it suggests simple pseudo-labeling of unlabeled image collections can suffice for learning transferable representations.
Why pretraining compensates for label noise is framed as a hypothesis that pretraining transfers general structural knowledge rather than details, and this is not directly verified. The plateau at the high-quality end and the apparent compensation of quality by dataset volume still need mechanistic characterization. Conclusions rest on abdominal CT with DynUNet, so whether other modalities, tasks, and architectures show the same pattern is an open question. In addition, MedSAM is deliberately used as a proxy for low-quality labels, and the artifacts from stacking its 2D predictions into 3D belong to that proxy construction rather than to a general claim about interactive foundation models.
