Skip to main content
Back to timeline
arXivSource publication:

Calibrating rater differences in prototype space with attention lets few-shot medical segmentation emit per-rater predictions and raise Dice on CURVAS and QUBIQ

Synopsis

The work formalizes few-shot multi-rater medical image segmentation and proposes a prototype-centric personalization framework: a consensus mask and consensus prototype are averaged from multi-rater masks, each rater prototype's deviation from the consensus prototype serves as the attention key, and a shared self-attention module calibrates the concatenated rater prototypes, combined with a calibration loss, pseudo-style supervision synthesized from superpixel pseudo labels via random structured boundary transformations, and two-stage training; on CURVAS abdominal CT (kidney, pancreas, liver; three experts; 20 training and 65 test scans) and QUBIQ brain-growth MRI (one class, seven raters; 34 training and 5 test scans), evaluated by per-rater Dice, the method improves consistently over pro

Source-provided article image: Attention-Based Prototype Calibration for Multi-rater Few-Shot Medical Image Segmentation
Fig. 1

Fig. 1: Overview of the proposed framework. (a) Few-shot multi-rater segmenta- tion pipeline. Rater-specific and consensus prototypes are extracted from support features, and calibrated by the Joint Attention-based Prototype Calibration (JAPC) module to generate R personalized predictions via cosine similarity with query features. (b) Internal structure of JAPC, modeling structured inter- actions between rater prototypes and their deviations. (c) Pseudo-style genera- tion. Pseudo labels derived from superpixel/supervoxel segmentation are trans- formed to synthesize multiple rater-style variants, which are used as support labels and query supervision during training. The “Network” block in (c) corre- sponds to the full pipeline shown in (a).

· Page 4

Interpretation

The paper formalizes few-shot multi-rater segmentation as a new problem: each support image carries R rater masks per class, and the model must output R personalized predictions per query, one per rater style observed in the support set, rather than a single consensus segmentation. Prior few-shot medical image segmentation methods typically assume a single reliable annotation per image, while multi-rater segmentation has mostly been studied in fully supervised settings with abundant annotations; this work places scarce supervision and heterogeneous annotation in one setting. The contribution is presented as a problem formulation and protocol definition, with explicit notation for support and query sets under R raters, developed within the existing N=K=1 few-shot protocol.

The paper proposes JAPC (Joint Attention-based Prototype Calibration): deviations between consensus and rater prototypes serve as keys, queries come from concatenated rater prototypes, and values are the prototypes themselves, yielding calibrated prototypes that are matched to query features by cosine similarity; the module leaves the backbone feature extractor unchanged and can be plugged into any prototype-based few-shot segmentation method. Relative to a naive baseline that processes each rater independently, this design lets each rater prototype attend to the deviations of all other raters, explicitly modeling shared semantics and rater-specific boundary tendencies instead of treating annotation differences as random noise. The paper gives the attention computation and the calibration loss Lcalib (mean squared difference between rater prototypes and calibrated prototypes), and the Abd-CT Setting 2 ablation shows attention and two-stage training add further gains, with all modules combined giving the best overall result.

The paper introduces pseudo-style generation and two-stage training: superpixel pseudo labels are transformed by random structured boundary operations such as dilation and erosion to synthesize R style variants used only during training for support labels and query supervision, and within a total budget of T iterations the base prototype model is trained for the first T/2 iterations before JAPC and pseudo-style are enabled. Superpixel pseudo labels are style-agnostic and lack rater-specific variation; pseudo-style generation supplies multi-style supervision while retaining pseudo supervision, and the two-stage schedule targets unstable prototype dynamics when the calibration module is trained with a randomly initialized encoder. The ablation shows pseudo-style supervision gives a modest improvement, attention and two-stage training further enhance performance, and calibration loss significantly strengthens prototype regularization; implementation details state 50k iterations on Abd-CT with the method initialized from the 25k checkpoint, and 10k iterations on Brain-MRI starting from 5k.

On two multi-rater benchmarks the method consistently outperforms few-shot and multi-rater baselines in per-rater Dice, with gains stable across annotators. The paper reports SSL-ALPNet rising from 61.48 to 62.79 and DSPNet from 61.52 to 62.25 on Abd-CT Setting 1, DSPNet from 54.02 to 66.62 on Brain-MRI, and in Setting 2 SSL-ALPNet from 56.58 to 58.69 and DSPNet from 58.11 to 59.62, with per-rater macro-averaged Dice rising in step. All few-shot methods use identical support-query splits and the same DeepLabv3-ResNet101 backbone, with a fully supervised model included as an upper bound; the paper also reports per-class, per-rater, and qualitative results, and notes that pancreas performance may still lag behind supervoxel-based methods that benefit from stronger region-level constraints.

Perspective

The work targets few-shot medical segmentation where the support set contains multiple rater masks and personalized per-rater predictions are required, applied to multi-organ abdominal CT and single-class brain MRI multi-rater data; because JAPC does not modify the backbone feature extractor, it can be layered directly onto existing prototype-based few-shot pipelines, offering later work a pluggable personalization component and a per-rater evaluation baseline. For readers, this means that in tasks where rater style differences are structured, modeling those differences explicitly in prototype space is worth trying over collapsing to consensus or training raters independently.

In the Abd-CT Setting 2 ablation, adding pseudo-style lowers pancreas Dice from 35.32 to 34.21 while the full module combination reaches 36.78, indicating that components act differently across organs and that readers should watch this class-level variation. The paper also notes pancreas performance may still lag behind supervoxel-based methods that benefit from stronger region-level constraints, and that pseudo-style supervision reduces fragmentation but does not always eliminate the gap. In addition, Brain-MRI is evaluated on the validation set because the official test set is unavailable, with only 5 test scans and a single class, so the reproducibility of its larger gain on bigger multi-rater cohorts remains an open question.

Sources