Skip to main content
Back to timeline
arXivSource publication:

SAGE suppresses prompt-invariant visual attention sinks to improve grounding and reduce hallucination across VLM families

Synopsis

Revisiting visual attention sinks in vision-language model decoders, this work finds a layer-dependent structure in which early and final decoder layers repeatedly collapse onto the same few image regions across prompts (Prompt-Invariant Sinks, PIS) while mid layers stay prompt-conditioned and drive vision-language alignment; it then proposes SAGE, a training-free intervention that applies an additive bias to visual-token attention logits using token-aligned ROI masks from standard vision backbones such as CLIP, ViT, and DINOv3, steering attention away from PIS toward query-relevant regions and improving visual grounding, reducing hallucination, and yielding gains on public benchmarks across LLaVA-1.5, Qwen2-VL, InternVL3, and NVILA.

Source-provided article image: SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders
Figure 1 ·

Figure 1: Prompt-Invariant Sinks and prompt-conditioned mid-layer attention in a VLM decoder. We visualize decoder attention over visual tokens for three different prompts (rows) at representative early, mid, and final decoder layers (columns). Early and final layers exhibit prompt-invariant attention collapse onto nearly identical image regions across prompts, forming Prompt-Invariant Sinks. In contrast, mid-layer attention is prompt-conditioned: it shifts to the menu board when asking for the espresso price (Prompt 1), to the wall clock when asking for the current time (Prompt 2), and becomes broader for holistic scene description (Prompt 3). Warmer colors indicate higher attention.

arXiv

Interpretation

The paper identifies Prompt-Invariant Sinks (PIS): across three substantially different prompts, early and final decoder layers repeatedly assign high attention to nearly identical image locations even though the question changes, whereas mid layers are the only consistently prompt-responsive part, shifting to the menu board for an 'espresso price' query and to the clock for a time query. Prior work established the existence of visual attention sinks but largely treated them as a uniform effect when attention is aggregated across layers; this study analyzes layer by layer and shows sink behavior is structured across depth rather than uniform. Evidence comes from layer-wise attention visualizations (Figure 1) and two independent diagnostics: cross-prompt IoU of the sink set and top-1 spectral concentration, which correlate strongly across LLaVA-1.5 and Qwen2-VL on TextVQA and GQA.

The paper proposes SAGE, a training-free inference-time intervention that applies an additive bias to visual-token attention logits at selected decoder layers, amplifying ROI tokens and suppressing non-ROI tokens to reallocate the attention budget while preserving the original model parameters. ROI masks can be produced by multiple standard vision backbones and mapped into token-aligned masks; CLIP conditions on the query while ViT and DINOv3 provide query-agnostic saliency, yet all three drive the same intervention rule. The method is a logit-level bias applied at inference with no gradient updates; it is instantiated on LLaVA-1.5, Qwen2-VL, InternVL3, and NVILA with a fixed bias strength, together with single-layer sweeps and multi-layer combinations.

Across public benchmarks, SAGE improves over the no-intervention baseline, with particularly strong gains on the MMVP family that stresses fine-grained pairwise visual discrimination, for example raising LLaVA 1.5-13B MMVP Pair from 11.85 to 37.78 with CLIP ROI. Gains appear at selected decoder depths, and the effective depth varies across model families and benchmarks; among multi-layer sets, mid and final compositions (Mid-final 62.60, Final-final 62.42) outperform very-early ones (Very-early only 19.19). Table 1 reports the best single-layer result per model, benchmark, and ROI source, which is an upper bound on the achievable effect under the layer sweep; Table 2 reports mean TextVQA accuracy grouped by normalized depth.

Fixed-layer and ROI-content controls show the gain depends on what the target region contains: on LLaVA-1.5-7B, with a single layer chosen in advance from label-free diagnostics, TextVQA rises from 44.0 to 54.6 and GQA from 57.1 to 62.3, while replacing the ROI with a randomly selected set of visual tokens of the same size leaves every benchmark at or below baseline. This control separates the effect of steering toward the right region from a generic perturbation of the attention distribution. The fixed-layer setting never re-selects the layer per benchmark, and the random-ROI control holds the intervention layer, bias strength, and mask size fixed while replacing only the mask content.

Perspective

The work targets vision-encoder plus decoder-only LLM VLMs and steers attention at inference with token-aligned ROI masks, suiting questions whose decisive evidence lies in a single localized region, such as reading text on a distant sign or verifying a printed number on an aircraft. For engineering readers who want to improve grounding and fine-grained discrimination without retraining, SAGE offers a reusable interface: masks can come from CLIP, ViT, or DINOv3, the intervention layer can be single or multi-layer, and the bias strength is fixed globally. The paper also reports fixed-protocol results across four model families, indicating the intervention is not confined to one architecture.

SAGE's effectiveness depends on ROI estimate quality: fragmented or misaligned masks can steer attention toward irrelevant tokens and suppress the true evidence, and the random-ROI control shows an uninformative mask removes the gain entirely, so a deployed system may need a confidence signal to decide when to intervene. The same mask is broadcast to all decoding steps and all attention heads, which suits questions with a single decisive region better than multi-region descriptions or multi-step reasoning. The diagnosis is at the level of depth rather than individual tokens, so whether explicitly suppressing sink tokens would add to the depth-level effect is untested, as is whether the same normalized depth transfers across model families without re-running the diagnostic. The paper does not explain why early and late layers develop prompt-invariant sinks while mid layers remain prompt-responsive, and treats this formation mechanism as an open problem.

Sources