Skip to main content
Back to timeline
arXivSource publication:

KAIST team turns prompt-template disagreement into preference supervision, lifting open-vocabulary segmentation across the MESS benchmark without pixel-level labels

Synopsis

The work proposes a preference-guided adaptation framework for open-vocabulary semantic segmentation that converts the segmentation differences different prompt templates produce on the same image (prompt disagreement) into binary preference supervision, adapting models with Region-Localized Preference Optimization (RLPO) plus consistency regularization in a streaming single-step setting, achieving consistent gains on the five domain groups of the MESS benchmark across SAN and CAT-Seg backbones (ViT-B/16 and ViT-L/14) without any pixel-level annotation and remaining effective under noisy preferences.

AI-generated editorial illustration: Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement

Interpretation

The authors observe that different prompt templates produce systematically different segmentations for the same image and class name, a phenomenon they call prompt disagreement, and repurpose it as a built-in source of preference supervision. Prior visual preference learning relies on auxiliary mechanisms such as stochastic sampling, input perturbation, or output sampling to create candidate diversity, whereas this work exploits variation already present in the prompting interface, adding no extra mechanism or hyperparameter. The paper reports that validation mIoU varies widely across templates for the same backbone in diverse specialized domains; in ablations, prompt disagreement reaches a mean mIoU of 45.88, above MC Dropout at 39.09 and test-time augmentation at 44.66 under a matched candidate size.

It introduces Region-Localized Preference Optimization (RLPO): cross-prompt entropy localizes high-uncertainty regions, the most disagreeing template pair is selected within each region, and the binary preference is applied to pixel-level segmentation scores, yielding dense spatial supervision from a single comparison. Standard DPO targets language generation, and prior segmentation preference work is largely limited to fixed-target or small closed label spaces; this work specializes the Bradley–Terry objective to template-conditional segmentation and uses a class-balanced regional score to keep large classes from dominating. On CAT-Seg-L, removing RLPO drops mean mIoU from 45.88 to 40.51, and removing consistency regularization drops it to 43.45, indicating RLPO is the primary driver of adaptation.

It adds consistency regularization that uses the preferred prediction as a pseudo-target for the rejected prediction outside the queried region, curbing unintended drift from preference optimization. The preference loss only supervises the selected region; this regularizer fills the gap of unconstrained predictions outside it, stabilizing single-step streaming updates. Ablation shows that without consistency regularization the mean mIoU is 43.45 versus 45.88 for the full method, and every domain group is lower than the full method.

On the five domain groups of the MESS benchmark, the method yields consistent gains across four backbone configurations without pixel-level annotation and remains effective under noisy preferences. Relative to the zero-shot baseline, mean mIoU improves by roughly 7.10, 6.55, 6.59, and 10.62 for SAN-B, CAT-Seg-B, SAN-L, and CAT-Seg-L, surpassing the matched-budget dense-mask reference on several specialized domains such as Medical Sciences and Engineering. Results are averaged over three seeds; in noise experiments the gain is essentially unchanged when labels are flipped with probability 0.05, and the method still recovers a substantial margin under heavier noise; Appendix D reports 0.920 overall human-oracle agreement over 600 judgments.

Perspective

The framework targets specialized-domain adaptation where the target vocabulary can be reasonably expressed through natural-language prompts and where prompt-induced candidates contain at least one useful hypothesis within the queried region; adaptation follows a streaming single-step protocol in which each image yields one binary preference and one gradient update, without revisiting images. It suits practitioners who want to transfer open-vocabulary segmentation models to medical imaging, remote sensing, industrial inspection, and agriculture at very low annotation cost (one binary judgment per image), and it has been validated on the five domain groups of the MESS benchmark with SAN and CAT-Seg backbones.

When all templates produce similarly inaccurate predictions within the queried region, the preference signal becomes less informative, since choosing the less wrong candidate offers limited guidance toward the correct segmentation; this is more likely for highly domain-specific concepts whose terminology, appearance, or label granularity is not well captured by natural-image-style templates, and the authors propose learned, domain-specific, or expert-provided templates as a future direction. In addition, main-experiment preferences come from an oracle that compares candidates against ground truth, validated in Appendix D by 600 human judgments (0.920 overall, 0.840 on hard pairs), so the behavior distribution of real annotators remains worth watching. This reading is full text, but figures and tables are rendered as text, so exact readings of individual numbers should be checked against the original tables.

Sources