Skip to main content
Back to timeline
arXivSource publication:

MLLM Hallucinations Arise When Information Distribution Drifts Inside Synergy Heads

Synopsis

The work proposes HEAL, which applies causal noise intervention to multi-head outputs and counterfactual Difference-in-Differences to categorize attention heads into redundant, visual, language, and synergy types, finds that hallucinations occur when the visual-language information distribution inside synergy heads drifts away from a healthy equilibrium rather than correlating strongly with the quantity or strength of modality-specific heads, and accordingly injects dynamic calibration factors into the value vectors of synergy heads at inference time, reducing hallucinations across multiple MLLMs and benchmarks.

AI-generated editorial illustration: MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

Interpretation

It proposes HEAL, which uses causal noise intervention and counterfactual Difference-in-Differences to categorize attention heads into redundant, visual, language, and synergy types, offering a head-level interpretable view. Whereas existing attention-based methods rely on indirect signals such as attention weights, HEAL directly intervenes on head outputs and measures changes in the layer representation, disentangling information composition at head level. The paper reports the taxonomy across several model families including LLaVA, Qwen, and InternVL, and reports 92.13%-95.36% head assignment agreement under alternative masking strategies such as zero masking, uniform masking, and swapping actual tokens.

It identifies visual-language information disequilibrium within synergy heads as a causally intervenable factor behind hallucination, and characterizes a task-driven phase transition of attention heads. Instead of attributing hallucination solely to macro-level modality heads, it locates the cause in the drift of information distribution inside synergy heads. Bidirectional causal analysis shows that artificially decreasing the visual information proportion can induce hallucination while increasing it can alleviate hallucination; the visual-language ratio difference between correct and hallucinated tokens in synergy heads remains observable across masking strategies and threshold settings.

It introduces a theoretically grounded dynamic calibration strategy that uses an equilibrium factor to regulate the visual-language information distribution and predictably steers the representation geometry toward equilibrium. Unlike attention enhancement that strengthens a single modality, HEAL performs an equilibrium reallocation between visual and language information, supported by two theorems stating that calibrating the value space is equivalent to calibrating the information distribution and that the equilibrium factor has a monotonic directional effect on visual alignment. As a plug-and-play module, it reports consistent improvements on hallucination and comprehensive metrics across POPE, CHAIR, MME, LLaVA-Bench, MMHal-Bench, and BLINK-Twice on LLaVA-1.5-7B, LLaVA-NeXT-7B, Qwen2-VL-7B, Qwen2.5-VL-7B, InternVL-7B, Qwen3-VL-8B, and InternVL3.5-8B.

Perspective

The result targets inference-time attention-head calibration and applies when both visual and language information are already present inside the model but the decoding-time distribution is imbalanced; the paper reports results across model families including LLaVA, Qwen, and InternVL and benchmarks including POPE, CHAIR, MME, LLaVA-Bench, MMHal-Bench, and BLINK-Twice, and notes that performance is stable when the equilibrium factor is within 0.4-0.6 and the update interval within 5-15 steps, which can serve as practical selection ranges.

The equilibrium factor and update interval are currently determined empirically, and the paper states that no closed-form theoretical rule is available, with optimal values varying across models and tasks; hallucinations caused by early visual encoding failures or a fundamental lack of visual evidence in the input cannot be resolved by attention-head calibration alone; the paper notes that existing benchmarks do not explicitly isolate rapid switching between visual description and language reasoning, making its impact hard to quantify directly; in addition, the loaded material is the full text, and if figure details are not fully rendered, conclusions tied to specific figures would still need confirmation against the original figures.

Sources