Splitting prompt dependence into prompt ambiguity and local sensitivity yields two low-correlated metrics that both correlate negatively with gynecological pelvic MRI segmentation quality
Synopsis
The work introduces the first framework that explicitly disentangles prompt dependence into prompt ambiguity (inter-user variability) and local sensitivity (interaction imprecision): a mixture density network models the image-conditioned distribution of plausible prompts and quantifies ambiguity by the trace of its covariance, while a stability margin captures the minimal perturbation under measurement noise needed to induce a mask change; evaluated on two female pelvic T2-weighted MRI datasets, UT-EndoMRI (81 patients with endometriosis, uterus segmentation) and MOGaMBO (94 patients with locally advanced cervical cancer, bladder segmentation), with MobileSAM and MedSAM, the two metrics correlate significantly and negatively with Dice, correlate little with each other, and conditional samp
Interpretation
It proposes a decomposition of prompt dependence into two computable components, prompt ambiguity and local sensitivity, via a variance decomposition based on the law of total variance: local sensitivity is the expected local variance induced by perturbations, and prompt ambiguity is the variance across reference prompts. Prior prompt-refinement methods (inferring an optimal prompt with detection networks, or iteratively generating and selecting prompts) and mask-level uncertainty methods (ensembles or stochastic inference) treat local sensitivity and prompt uncertainty together; this work separates them at the formulation level for the first time. The derivation applies the law of total variance to a perturbation variable and uses a first-order expansion for prompt ambiguity, giving VarB[f(B)] ≈ ∇bf(μB)ᵀΣB∇bf(μB); this is an analytical argument that does not rely on new data.
It approximates the image-conditioned distribution of plausible prompts pθ(b|x) with a mixture density network, quantifies prompt ambiguity as the trace of the predictive covariance Uamb(I), and further samples prompts to produce pixel-level entropy uncertainty maps. It turns prompt ambiguity from a qualitative notion into an estimable image-level scalar and propagates that distribution to pixel-level uncertainty, rather than producing only a single prompt or a single mask-level uncertainty. The MDN is a two-layer MLP with K=8 Gaussian components (K chosen by validation-set BIC), trained with negative log-likelihood, reaching validation nll ≈ −2.5; in uncertainty-map evaluation, conditional sampling (CondMDN) with MobileSAM reaches AUROC 0.98 and ΔEntropyok/err 0.45 (UT-EndoMRI) and 0.38 (MOGaMBO), above jittered prompts (AUROC 0.93) and a non-conditional distribution D0 (AUROC 0.96).
It proposes the stability margin δ⋆(I,b) as a measure of local sensitivity: at tolerance τ=0.1, the smallest perturbation amplitude whose average mask difference exceeds the threshold, approximated by a pretrained stability-margin predictor. Instead of estimating local sensitivity from infinitesimal perturbations, it gives a thresholded definition tied to a clinically acceptable discrepancy, and the empirical procedure requires no additional ground-truth masks, relying only on the model's response to synthetic perturbations. The predictor is a two-layer MLP trained with ℓ1 loss, reaching validation ℓ1 ≈ 0.001; offline oracle margins use 32 samples per perturbation type (translation or scale) over a discrete grid δ∈{0.005, 0.01, 0.05}.
Across two datasets and two models, prompt ambiguity and local sensitivity correlate little with each other, while each correlates significantly and negatively with Dice, indicating that they capture complementary failure modes. This provides empirical support for the disentangled design: if the two were highly correlated, the split would be unnecessary; low correlation indicates they capture different sources of prompt-related failure. With MobileSAM, r(Uamb, δ⋆) = −0.07 (MOGaMBO) and 0.001 (UT-EndoMRI); r(Dice, Uamb) = −0.45 (MOGaMBO) and −0.30 (UT-EndoMRI), r(Dice, δ⋆) = 0.36 and 0.44, all significant at p<0.001; the high-ambiguity, high-sensitivity class has Dice+/+ = 0.29 versus Dice−/− = 0.73 (MOGaMBO, MobileSAM).
Perspective
The framework targets medical imaging workflows that use promptable segmentation models and where prompts come from different users, especially organ delineation settings with inter-observer variability; it needs only a frozen image encoder and the model's responses, not additional ground-truth masks, making it suitable as a pre-deployment reliability assessment and failure-mode diagnostic tool. The authors state that next steps include integrating this evaluation framework into clinical practice, improving target delineation in radiotherapy, and extracting robust organ and lesion biomarkers; the code is publicly available.
The authors explicitly note that further experiments are needed to understand the impact of hyperparameters, starting with the number of components used to model pθ(b|x) and the stability margin threshold τ. In addition, the main driver of prompt dependence varies across datasets, models, and metrics: in MOGaMBO, Dice is more strongly anti-correlated with Uamb, whereas in UT-EndoMRI, Dice is more strongly anti-correlated with δ⋆, suggesting the driver is dataset-specific and differently encoded by image encoders. Readers should also note that current evidence comes from two female pelvic MRI datasets and two SAM-based models, so generalization to other organs, modalities, and models remains an open question.
