RPG-SAM pairs reliability-weighted prototypes with geometric adaptive thresholds to lift training-free one-shot polyp segmentation by 5.56% mIoU on Kvasir
Synopsis
The work proposes RPG-SAM, a SAM2-based training-free one-shot polyp segmentation framework that uses Reliability-Weighted Prototype Mining (RWPM) to weight support foreground prototypes by a contrast factor and a reverse purity factor while using background prototypes as negative anchors for noise suppression, Geometric Adaptive Selection (GAS) to pick binarization thresholds dynamically from morphological solidity and scale consensus, and a Prior-guided Iterative Refinement (PIR) loop to polish boundaries, reaching 78.65% mIoU and 85.65% mDice on Kvasir, surpassing ProtoSAM by 5.56% and 4.11% respectively, with comparisons also reported on three-center PolypGen, CVC-ClinicDB, and CVC-ColonDB.
Fig. 1: Motivation of RPG-SAM. (a) Regional and Contextual Heterogeneity. Top: The varying regional utility of the support foreground means that de- graded areas (e.g., reflection d) trigger false-positive noise in query images. Bot- tom: Treating the contextual background as a distinct information layer provides negative anchors to effectively suppress heatmap activations. (b) Intensity Het- erogeneity. The histogram of optimal binarization thresholds across the Kvasir dataset [8] highlights the stochastic response intensities of different query scenar- ios, exposing the limitations of fixed-threshold (homogeneous) sampling rules.
· Page 2Interpretation
Introduces RWPM, which weights SLIC superpixel foreground prototypes by the product Wk of a contrast factor Ck and a reverse purity factor Rk, and subtracts background prototypes as negative anchors from the heatmap to suppress false positives triggered by degraded regions such as specular reflections. Prior matching-based prompt sampling pipelines such as PerSAM, Matcher, OPSAM, and ProtoSAM treat support foreground pixels as equally representative and largely neglect the support background as a distinct information layer; this work explicitly separates prototype reliability from background contrastive reference. Ablation shows background suppression alone yields a 3.78% mDice gain on Kvasir, and adding RWPM raises mIoU from 67.89% to 71.27% and mDice from 78.72% to 81.05%.
Introduces GAS, which scans a range of thresholds to generate candidate masks and selects the best prior mask by a geometric score Sgeo combining weighted solidity and scale consensus, replacing a fixed threshold. Fixed-threshold or regional-statistics binarization rules cannot adapt to the stochastic response intensities across query images; GAS recalibrates the threshold dynamically using morphological priors. In ablation, GAS improves mDice by 2.59% over the best fixed threshold (τ=0.7); hyperparameter analysis shows the scan range [0.4, 0.7] gives the optimal candidate pool, with lower ranges introducing noise and higher ranges limiting diversity.
Introduces PIR, which uses the geometric prior mask as reference, diagnoses error type via coverage and IoU, samples the geometric center of the false-negative region via Euclidean Distance Transform as a positive prompt and the false-positive region as a negative prompt, and iteratively calls SAM2 to refine boundaries. It wires SAM2's boundary-refinement capability into an automated loop without manual intervention, and selects the mask with the highest IoU relative to the prior across the iteration history as the final prediction. In ablation, adding PIR on top of BG Supp., RWPM, and GAS raises Kvasir mIoU from 75.25% to 78.65% and mDice from 83.64% to 85.65%; PIR thresholds τcov=0.9 and τiou=0.8 give the best results.
Validates overall performance on four public datasets: Kvasir at 78.65% mIoU / 85.65% mDice, the three PolypGen centers at 59.45/64.78, 61.04/67.14, and 76.75/83.56, CVC-ClinicDB at 70.18/77.76, and CVC-ColonDB at 61.16/68.77. Compared with SEGIC, PerSAM, Matcher, OPSAM, and ProtoSAM under identical support-query pairs and backbone settings, RPG-SAM reports the highest values on each dataset shown. All inference runs on a single NVIDIA RTX 3090 with no training or fine-tuning; DINOv2 ViT-L/14 (560×560) serves as the feature encoder and SAM2 (1024×1024) as the mask generator.
Perspective
The framework targets label-scarce clinical settings, driving query-image segmentation from a single labeled support image, and applies to endoscopic polyp images; it builds on DINOv2 features and the SAM2 mask generator, with all inference on a single RTX 3090 and no training or fine-tuning. The authors state that future work will extend the framework to exploit temporal consistency in endoscopic video, indicating that the current results concern the static-image setting.
The ablation and hyperparameter analysis are concentrated on the single Kvasir dataset, while other datasets report only final comparison numbers, so each module's separate contribution in cross-center settings remains for the reader to judge. GAS relies on a reference area Aref representing the expected median polyp scale, and how that quantity is set is not expanded in the text. In addition, the code is marked as to be released, so implementation details cannot currently be checked; readers needing reproduction should watch this open item.
