Skip to main content
Back to timeline
arXivSource publication:

Using OCR Heads to Verbalize Image Semantics: General Verbalization Heads and a Verbalization Lens in VLMs

Synopsis

Across four vision-language models (Qwen3-VL-2B, Qwen3-VL-8B, Molmo2-7B, Llava-Next-34B), the work identifies attention heads causally necessary for OCR and shows they are general-purpose verbalization heads; summing their output-value matrices yields a verbalization lens that decodes image-token hidden states into interpretable semantic words at any layer, including layer 0, and whose pseudo-inverse provides concept vectors for editing image representations (e.g., replacing a tractor with a revolver), supporting the view that image representations are aligned with language space from early layers.

Source-provided article image: Using OCR Heads to Verbalize Image Semantics
Figure 1 ·

Figure 1 : We find that VLM attention heads responsible for OCR are actually generic verbalization heads that can read out semantics of image tokens that do not contain text. (a) Verbalization heads attend to text when prompted to transcribe words, but manually redirecting their attention to image tokens that do not contain text causes the model to output words corresponding to the semantics of those tokens (top 10% of Qwen3-VL-8B verbalization heads). (b) We repurpose these heads’ attention weights to create a verbalization lens that can be applied to image representations from any layer. Here, our approach reveals alignment of image representations with language space starting from layer 0, which logit lens is unable to see.

arXiv

Interpretation

It identifies attention heads causally necessary for OCR and shows they are not OCR-specific but general verbalization heads that output semantic words for non-text image tokens. Prior work on OCR mechanisms in VLMs largely stayed at attention visualization or behavior; here each head is scored with logit lens (Equation 1) and causally tested by mean ablation: ablating the top 10% of heads by OCR score fully degrades task performance, while ablating an equal number of random late-layer heads is less effective; forcing attention onto non-text tokens such as a bird wing or a branch makes Qwen3-VL-8B output _feathers or _branch. Consistent across four models and three model families; OCR scoring uses n=1024, ablations run on 100 images across 10 seeds; models reach 0.80-0.93 OCR accuracy on ImageNet backgrounds (only 0.05-0.75 on white backgrounds, hence realistic backgrounds).

Summing the verbalization heads' OV matrices into a single linear transformation (Equation 3) and combining it with vocabulary projection yields a verbalization lens (Equation 4) that reads semantic labels from image tokens at all layers, including layer 0. Raw logit lens is noisy in early-middle layers, whereas this lens gives interpretable results from layer 0, indicating that image representations are already aligned with language space when passed to the language decoder; on COCO single-token categories (56/80 for Qwen3 models), the gap in P(o) between images containing and not containing an object is larger than for raw logit lens and consistent across layers, with near-ceiling ROC-AUC. On COCO validation, n=512 images with the object and n=512 without per category (n=128 for Llava-Next-34B); compared against raw logit lens and an equal number of random heads; using the same LLM-judge metric as LatentLens, it improves Llava-Next-34B by 40.5 points on average and is on par for other models.

The subspace read by the verbalization lens is causally relevant: its pseudo-inverse (Equations 5-7) yields concept latent vectors that can be added and subtracted across all layers and image token positions (Equations 8-9) to replace object concepts in an image. This goes beyond interpretability readout to give causal evidence that the subspace matters for the model's image descriptions; replacing an ant with an iPod preserves the original object's color and texture (described as a glossy, reddish-brown classic model), and replacing a tractor with a revolver leads the model to synthesize wheels, license plate, and driver into a surreal revolver creatively modified to resemble a classic car. LLM-judged (o4-mini) on n=256 random ImageNet images along prevalence of added/removed concepts, specificity, and coherence; edits require alpha>1, while alpha=10 damages specificity; uses the rank explaining the top rho=60% of L_p energy, with rank and alpha swept on 100 images first.

Verbalization heads partially overlap with previously studied head types, suggesting OCR as a narrow task can serve as an entry point for identifying more general pixel-to-semantics mechanisms. Six of the 11 filter heads found in Qwen3-1.7B fall in the top 10% of verbalization heads; 30% (3/10) of gaze heads in Qwen3-VL-2B and 29% (29/100) in Qwen3-VL-8B are also verbalization heads, above random-sampling expectation, indicating a relationship rather than identity. Overlap statistics against published head sets with random-overlap expectations as a control; this is correlational evidence, and the authors frame it as a possible relationship.

Perspective

The results apply to the four studied VLMs (Qwen3-VL-2B, Qwen3-VL-8B, Molmo2-7B, Llava-Next-34B) and their image-token representations; the verbalization lens targets settings that need fast decoding of hidden states into semantic labels, and the editing experiments focus on single-token concepts validated on Qwen3-VL models. For readers wanting to reuse the method, the public code and interactive demo support cross-model readout; for those interested in multi-token concepts, larger concept sets, or pixel-level localization, this setup is a starting point rather than an endpoint.

Readers may still watch: token-level localization mIoU and mAP are overall low (e.g., 0.190/0.407 for Qwen3-VL-8B), which the authors partly attribute to register token behavior, and how broadly that explanation holds remains open; edit effectiveness depends on sweeping rank and scaling factor alpha, with large alpha (e.g., 10) hurting specificity and different models needing different alpha; concept editing is currently limited to single-token concepts, leaving multi-token or more abstract concepts to be observed; and the overlap with filter heads and gaze heads is correlational, so the exact functional boundary of verbalization heads remains to be characterized.

Sources