SCISSR swaps point and box prompts for scribbles, reaching 95.41% Dice on EndoVis 2018 and 96.30% Dice on cross-domain CholecSeg8k
Synopsis
The work presents SCISSR, a scribble-promptable framework for interactive surgical scene segmentation: a lightweight Scribble Encoder turns freehand scribbles into dense prompt embeddings compatible with the mask decoder, and together with Spatial Gated Fusion and toggleable LoRA adapters it supports multi-round correction over a frozen SAM 2 backbone, reaching 95.41% Dice on EndoVis 2018 with five interaction rounds and 96.30% Dice on the unseen CholecSeg8k with three rounds, outperforming iterative point prompting on both benchmarks.
Interpretation
A lightweight Scribble Encoder maps freehand scribbles to dense prompt embeddings, letting scribbles plug directly into the mask decoder and support round-by-round correction. Prior interactive methods are predominantly click-based, and most surgical adaptations of SAM/SAM 2 retain point or box prompting; ScribblePrompt targets intensity-normalized biomedical images and does not use SAM 2's memory bank, while other works use scribbles only as weak supervision. The encoder is specified structurally (resize to 256x256, two stride-2 convolution blocks plus a 1x1 projection to D=256, producing a 256x64x64 embedding) and evaluated on EndoVis 2018 and CholecSeg8k under a fixed-round automated protocol.
A dual-track scribble pathway with memory-driven iterative refinement: Track 1 accumulates all scribbles as a dense prompt, while Track 2 injects only the latest correction into the Memory Attention query via Spatial Gated Fusion, with a memory bank holding just the previous round's features. SAM 2's temporal memory, originally for video, is repurposed for multi-round correction on a single image, and a zero-initialized learnable scalar alpha makes SGF an identity mapping at the start of training. Ablation on EndoVis 2018 shows R2 mIoU rising from 80.21% (baseline) to 83.59% with SGF and to 88.30% with SGF plus Memory; per-class ablation reports a +14.20 R2 IoU gain for Wrist.
An architecture-agnostic design: the Scribble Encoder, SGF, and LoRA adapters interact with the backbone only through standard embedding interfaces, so they transfer to prompt-driven architectures such as SAM 3 without structural modification, and the LoRA adapters are toggleable to preserve standard video propagation. Compared with adaptations tied to a single backbone, the new capability is confined to lightweight add-on modules while the backbone stays frozen. The paper instantiates on SAM 2 Tiny with LoRA rank r=8 and scaling alpha/r=2 inserted into query and value projections of the mask decoder and Memory Attention; the image encoder remains entirely frozen. Transfer to SAM 3 is stated as a design property, with no SAM 3 measurements reported.
Validation on one in-distribution and one out-of-distribution laparoscopic dataset shows scribble prompting beating point and box: contour reaches 91.60% mIoU / 95.41% mDice at R4 on EndoVis 2018, and adaptive reaches 92.30% mIoU / 95.82% mDice at R2 on CholecSeg8k. Point-prompt baselines often degrade as rounds increase (SAM3 1pt/CC on CholecSeg8k drops from 56.74% mIoU at R0 to 52.62% at R2), whereas scribble strategies improve steadily; box and supervised baselines do not support iterative refinement. EndoVis 2018 has 4,616 test samples and CholecSeg8k 8,800; on convergence efficiency SCISSR reaches Dice>=0.75 with 99.0% success in a mean 1.37 rounds versus 75.5% / 1.51 for the point baseline; a point-density study shows both SAM 2 Tiny and SAM 3 peak at 10 points per channel and degrade sharply at 30 and 50.
Perspective
The result targets interactive annotation and segmentation of laparoscopic surgical scenes, in settings built on a SAM 2-style prompt-driven segmentation model where a user can draw corrective strokes round by round; the authors state that future work will extend SCISSR to video-level annotation and validate usability with human annotators. Training uses EndoVis 2018 while CholecSeg8k is held out for out-of-distribution testing, so cross-domain conclusions are limited to the procedures and label sets covered by these two datasets.
Evaluation uses an automated protocol that synthesizes initial scribbles and corrective strokes from ground-truth masks rather than human operation, so real annotation efficiency and user experience remain to be validated. Scribbles are synthesized with four strategies (centerline, wave skeleton, contour, line); contour performs best across rounds while centerline lags, indicating sensitivity to scribble shape. The point-density analysis covers only a 4-video subset of CholecSeg8k. The paper mentions transfer to SAM 3 but reports no SAM 3 measurements, and Cystic Duct has only 1 sample in CholecSeg8k, so its per-class numbers are of limited stability.
