Skip to main content
Back to timeline
Intelligent ComputingSource publication:

Prompt Engineering in the Segment Anything Model: Methodologies, Applications, and Emerging Challenges

Synopsis

This survey systematically reviews prompt engineering research for SAM and its growing ecosystem, proposing a hierarchical taxonomy that organizes methods into geometric prompts, textual semantic prompts, and multimodal fusion prompts, further tracing the transition from manually crafted prompts to automated generation based on detector outputs, prototype learning, reinforcement learning, and vision-language models, while tracing how prompt engineering enables cross-domain generalization in medical imaging, remote sensing, industrial inspection, and anomaly detection, and identifying key challenges such as prompt sensitivity, cross-modal misalignment, and computational inefficiency alongside future directions including causal prompt reasoning, collaborative multi-agent prompting, and diffu

Source-provided article image: Prompt Engineering in the Segment Anything Model: Methodologies, Applications, and Emerging Challenges
Figure 1

Figure 1: Taxonomy of prompt types and generation methods.

· Page 4

Interpretation

Proposes a function-oriented hierarchical taxonomy of prompt engineering that categorizes SAM-related methods into geometric prompts, textual semantic prompts, and multimodal fusion prompts. Prior SAM surveys largely focus on architecture and downstream performance, or cover only a single prompt type such as point prompts; this work uses the functional question of what to segment and where to segment as its classification logic, offering a unified taxonomy across strategies. A survey work whose classification rests on systematic retrieval and synthesis of literature first made publicly available from SAM's April 2023 release through December 2025, drawn from Google Scholar, Semantic Scholar, arXiv, IEEE Xplore, and the ACM Digital Library.

Characterizes four progressive paradigms of geometric prompting: manual input and heuristic enhancement, detector-based automatic generation, self-prompting from image features, and optimization-driven generation. Organizes scattered automated prompting methods into a progressive spectrum along the nature of generation rules and the complexity of inference mechanisms, with a comparison across prompt rule, core driver, and interaction at inference. Based on synthesis and comparison tables of representative methods, covering heuristic (e.g., SAMAUG, RoBox-SAM), detector-based (e.g., YOLO-SAM 2, AM-SAM), self-prompting (e.g., APSeg, Label Anything), and optimization-driven (e.g., TEPO, SAMRefiner) approaches.

Delineates two technical routes for textual semantic prompting: semantic granularity-based prompt construction and text-to-geometric prompt generation. Distinguishes deep semantic injection from explicit surrogate coordinate transformation, noting that the former prioritizes embedding alignment while the latter emphasizes compatibility with SAM's native geometric interface. A survey synthesis covering category-level and part-level prompts (e.g., Talk2SAM, SP-SAM), structured-knowledge driven approaches (e.g., PG-SAM, TP-DRSeg), and CLIP-based alignment pipelines (e.g., CLISC, GenSAM).

Summarizes domain adaptation logic and quantitative benchmarking of prompt engineering in medical imaging, remote sensing, and industrial anomaly detection. Organizes cross-domain applications by task-driven paradigms (such as anatomical, pathological, and clinically-driven adaptive learning in medicine) and aggregates reported results across metrics including Dice, NSD, HD95, F1, and AUROC. A survey aggregation in which quantitative results come from each cited method's own reported benchmarks; the text also notes that differing metric choices across papers make direct numerical comparison challenging.

Perspective

The survey's scope covers methods made publicly available from SAM's April 2023 release through December 2025 that use prompts to steer and customize segmentation outcomes, excluding task-specific models lacking prompt-driven interaction; its taxonomy and comparisons are intended for researchers and practitioners seeking to understand prompt engineering's role in segmentation foundation models, particularly for adaptation design in specialized domains such as medical imaging, remote sensing, and industrial inspection.

As a survey, its quantitative comparisons rely on each cited method's self-selected benchmarks, and differing metric choices make cross-paper numerical comparison challenging; the text notes that systematic modeling of how prompts modulate attention distributions within SAM and ultimately influence the mask generation pathway remains lacking; directions such as causal prompt reasoning, collaborative multi-agent prompting, and diffusion-based progressive refinement are largely outlooks whose feasibility and effectiveness await further validation.

Sources