Skip to main content
Back to timeline
arXivSource publication:

Across 25 models, a shared L0 attention routing circuit marks VQ-tokenized VLMs, and ablating it alone cuts open-ended object hallucination by 31% relative

Synopsis

Using activation patching across twenty-five models spanning eight LLM families, the work identifies an early-layer (L0) attention routing circuit shared by VQ-tokenized vision-language models, proposes a three-gate diagnostic that isolates ten models carrying it and rejects the other fifteen, shows via a single-variable architectural swap (LLaVA-1.6 CLIP+MLP to VQ+Linear) that vector quantization is the source of the pathological signal, and finds that only L0 ablation reduces object hallucination in open-ended generation (CHAIR_i down 31% relative) while tuned DoLA and VCD do not.

Source-provided article image: Where Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models
Figure 1 ·

Figure 1: Methodology overview. Clean and corrupted image pairs are constructed and passed through the transformer; residual divergence and early-layer attention mass are measured to detect a routing circuit via a three-gate diagnostic. Models passing all three gates have a VQ-installed L 0 L_{0} circuit confirmed by a single-variable causal control: swapping only the projector type from MLP to VQ installs the circuit; a matched-compute MLP control does not. Ablating L 0 L_{0} reduces object hallucination in open-ended generation (CHAIR i drops 31% relative) while tuned DoLA achieves better binary calibration.

arXiv

Interpretation

The study identifies a cross-architecture early-layer (L0) attention routing circuit present in VQ-tokenized unified VLMs and derives a three-gate diagnostic that labels ten models positive (five natural unified-VQ VLMs across three LLM families plus five induced variants) and rejects the remaining fifteen. Prior decoding-time fixes treated object hallucination as generic miscalibration without an architectural account; this work ties it to specific architecture and pretraining properties and offers a diagnostic separating carriers from non-carriers. Activation patching across twenty-five models and eight LLM families, with an explicit positive-versus-negative split of ten against fifteen.

A single-variable architectural swap (LLaVA-1.6 CLIP+MLP to VQ+Linear) installs the circuit, whereas a matched-compute MLP control on identical data does not, isolating vector quantization as the source of the pathological signal; the routing pathway carrying it is one the backbone already provides. This is a controlled isolation of causal source rather than a correlational observation, pointing to vector quantization rather than general vision-encoder differences as the key variable. A single-variable swap paired with a matched-compute MLP control on identical data.

In open-ended generation only L0 ablation reduces object hallucination, with CHAIR_i falling 31% relative, while tuned DoLA and VCD leave it unchanged or worsen it; on binary calibration, tuned DoLA outperforms the baselines. Mechanism-agnostic decoding cannot replicate this targeted intervention, and improved calibration is shown to be distinct from reduced open-ended hallucination. Comparison against tuned VCD and DoLA baselines, reporting a 31% relative CHAIR_i reduction.

Perspective

The results target unified VLMs that tokenize images through a vector-quantized codebook, in grounded yes/no benchmarks and open-ended generation evaluation settings; for such models, L0 ablation is an actionable intervention point and the three-gate diagnostic can indicate whether a model carries the circuit. The paper does not draw applicability conclusions for models using other visual tokenization schemes or non-unified architectures.

A careful reader would still ask whether the three-gate diagnostic remains stable on new models beyond the twenty-five studied, whether L0 ablation keeps its benefit outside open-ended generation, and how much of the effect is attributable to vector quantization versus pretraining. The text is presented as an abstract without dataset, prompt-setting, or statistical detail, which are open questions to check when judging further.

Sources