Skip to main content
Back to timeline
arXivSource publication:

AnswerMap builds VLM spatial rationales from the outer product of row and column "yes" posteriors, and deleting its region flips 53% of correct answers across four models

Synopsis

The work introduces AnswerMap: the image is cut into K row and K column bands, each shown alone to a frozen VLM with a yes/no question about the query, and the outer product of the row and column "yes" posteriors gives a query-conditioned spatial map on which a fixed read-out (e.g., expectation, maximum) derives continuous outputs natively; across four models and three query distributions the map agrees with the model's own generated point at AUC 0.85 versus 0.38 for attention, deleting its region flips 53% of correct answers versus 19% for attention, and its maximum flags hallucinated objects, its expectation localizes when the model's own pointing fails, and its top-mass crop fixes about half of the model's wrong answers.

AI-generated editorial illustration: AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors

Interpretation

It introduces AnswerMap, a training-free, task-agnostic, black-box spatial rationale: the image is cut into row and column bands, each band is shown alone with the query as a yes/no relevance question to a frozen model, and the outer product of the row and column "yes" posteriors forms the query-conditioned spatial map. Unlike existing tools that rely on text rationales (mismatched modality, no explicit region testable against the image) or internal read-outs (originating before the answer is generated and requiring white-box access), each cell score here is a full-contrast recognition judgment of a region on its own rather than the marginal effect of removing that cell from the full image. The paper gives the method definition and cost: a K×K map costs 2K forward passes against 65 for occlusion at the same grid, and it needs only the first-token logprobs of two labels, so it runs unchanged on a closed model API (reported on GPT-6-sol).

It validates faithfulness with two ground-truth-free tests: the map should agree with the model's own generated point, and deleting the map's region should break a correct answer. The paper proposes these two tests as a protocol applicable to any spatial explanation of any VLM, and reports results across four models (Qwen3-VL-4B/8B/30B-A3B, InternVL3.5-8B) and three query distributions (RefCOCOg, RefCOCO+, CAVE). Agreement against the model's own point: AUC 0.85 versus 0.38 for attention; deletion flips 53% of correct answers versus 19% for attention and 8.2% for a uniform random region; a blank-image control lands exactly at chance on both metrics, and a swapped-image control shows the map follows the image it was shown.

The same map with different fixed read-outs serves different downstream tasks: the maximum flags hallucinated objects without generation, the expectation localizes correctly when the model's own pointing fails, and the top-mass region fed back as a crop fixes about half of the model's wrong answers. This treats the rationale as an output interface as well: for continuous-output tasks such as location it bypasses reliance on discrete text tokens, and image dependence is guaranteed by construction. On POPE the model says yes 4,006 times, 223 of them hallucinated, and the map maximum separates the two claim types at ROC-AUC 0.833; on the weaker pointer Lingshu-7B the expectation beats the model's own point by 11 points; on TextVQA the map's crop turns 50% of the model's 42 wrong answers right versus 38% for a random crop.

Multigrid products and ablations give usage guidance: composing coprime grids stays flat across a five-fold query budget, and product fusion beats mean and max. The paper contrasts refining one grid against composing coprime grids as two ways to spend a query budget, and notes that nested grids double-count correlated evidence. Ablations use Qwen3-VL-4B on RefCOCOg: refining one grid peaks narrowly at 12 queries and collapses as bands thin, while coprime composition reaches the same peak at 16 queries and stays flat; on fusion, product scores 1.50 versus mean 1.39 and max 0.89.

Perspective

The result is aimed at readers who need to check a VLM's spatial basis, such as a radiologist reading a generated report or an auditor checking a claim, and at developers who want to audit a model without weight access. The operator applies to models that accept non-square inputs, natively or by tiling, since each band is a long thin strip; the paper reports it runs with no per-model tuning on four open-weight models and one closed API model. The three read-outs map to different settings: the maximum for hallucination auditing of object-existence claims, the expectation for localization when the model's own pointing fails, and the top-mass crop for giving the model a close-up before it answers. The paper also proposes richer read-outs, training on top of the map as label-free supervision, and probe-guided generation as next steps.

The paper states two scope limits: the operator needs a model that accepts non-square inputs, and because the map is built from row and column answers it knows which rows and which columns contain the query but not which pairs of them do, so when a query matches two separate objects the map also lights the two empty cells where their rows and columns cross. The read-out is described as the main design surface, with richer read-outs (extent, count, relations) left open; the paper reports that the maximum reads spatial commitment rather than answer reliance on the image (AUC 0.49 against the whole-image label), so different read-outs answer different questions. This summary draws on the paper's full text and abstract and adds no data or experiments the paper does not report.

Sources