Purdue team's training-free one-shot segmentation framework lifts mean IoU by 33.36 points over the strongest baseline on pool-boiling images
Synopsis
The work presents a training-free, one-shot framework for scientific image segmentation that uses a single annotated reference image to combine frozen DINOv3 with SAM: a background direction estimated from the reference background drives background-adaptive orthogonal projection of DINOv3 patch features to suppress artifact-related directions, cosine similarity then localizes candidate regions, and generated prompts feed SAM. Evaluated on red-blood-cell microscopy, structured-illumination pool boiling, and chest radiography, it improves mean IoU over the strongest baseline by 5.15 points on microscopy and 33.36 points on pool boiling, while matching SAM2 on chest radiographs.
Figure 1 : Generalist-to-specialist one-shot segmentation across scientific imaging domains. A single annotated reference defines each target concept, and the same frozen DINOv3–SAM framework localizes and segments corresponding structures without task-specific training.
arXivInterpretation
A training-free, one-shot generalist-to-specialist framework that adapts frozen vision foundation models to a specific scientific segmentation task using only a single annotated reference image, avoiding supervised training per experimental setting. Compared with task-specific supervised training that needs extensive expert annotation, and with existing training-free one-shot approaches (GF-SAM, INSID3), the framework reduces adaptation cost to one reference sample while keeping foundation-model weights frozen. The method explicitly uses frozen DINOv3 and SAM, with the reference image and its mask defining the target structures; evaluation reports mean IoU and Dice across three datasets against SAM2, GF-SAM, and INSID3.
A background-adaptive orthogonal projection: the reference mask is downsampled to the DINOv3 patch grid, reference tokens are partitioned into foreground and background, a unit background direction is formed by averaging and L2-normalizing background tokens, and both the foreground prototype and target tokens are projected onto its orthogonal complement to suppress feature directions tied to background structures and imaging artifacts. Prior approaches such as INSID3 use a generic Gaussian-noise direction for orthogonal projection and do not adapt feature correction to the background and artifacts actually observed in the annotated reference; this work estimates the direction from that reference background. Across 20 pool-boiling frames, mean background similarity is reported to drop to near zero from higher levels for unprojected features and the Gaussian-noise projection, with a corresponding increase in contrast-to-noise ratio; Fig. 4 shows unprojected and Gaussian-noise projections retaining elevated responses over the structured background.
Validation of the same framework across three modalities, covering distinct target structures, artifacts, and background characteristics. Evaluation spans microscopy, structured-illumination reconstruction, and projection imaging rather than a single dataset. Red-blood-cell microscopy uses the 100k-RBC-PathOlOgics dataset (240,790 segmented RBC images across nine classes, with a separate annotated reference per class); pool boiling uses 20 structured-illumination reconstructed frames with SHT and RMS reconstructions; chest radiography uses the Montgomery dataset (138 frontal radiographs with lung masks, including 58 tuberculosis-positive cases).
Segmentation results show the largest gains when artifacts dominate learned feature representations: 33.36 points mean IoU and 26.62 points Dice over INSID3 on pool boiling, 5.15 points IoU and 2.94 points Dice over GF-SAM on RBC microscopy, and performance comparable to SAM2 on chest radiography. Relative to baselines on pool boiling, where residual illumination patterns, spatially varying contrast, and closely spaced features are present, the improvement concentrates on artifact-dominated modalities. Table 1 reports mean IoU and Dice for SAM2, GF-SAM, INSID3, and the proposed method on each dataset; Fig. 3 shows SAM2 and GF-SAM recovering only a subset of bubbles, INSID3 producing enlarged and merged masks, and the proposed method recovering more instances while better preserving shapes and boundaries.
Perspective
The framework targets scientific image segmentation where a single annotated reference must transfer to corresponding target structures, and it is demonstrated on microscopy, structured-illumination reconstruction, and projection imaging. Its gains are most pronounced when artifacts dominate learned feature representations, as in pool-boiling images. For researchers and practitioners seeking to avoid per-task training and large expert-annotated datasets, it offers a path to adaptation from one reference image. The reference mask need not contain every target instance, provided the annotated regions adequately represent the target structure.
Two improvement values are missing from the abstract and should be taken from Table 1 and Section 3.2 percentages; the specific background-similarity and contrast-to-noise-ratio values in Section 3.3 are not fully given in the text, only described in words. Performance on chest radiography is comparable to SAM2, indicating no additional gain on low-contrast projection images with overlapping anatomy, so the applicability boundary still needs observation across more modalities and artifact types. How the choice of reference image affects results, and behavior when the reference mask is not representative, are open questions a careful reader would watch.
