Skip to main content
Back to timeline
arXivSource publication:

SECRET intercepts cross-modal interference at the question relay, lifting audio-visual LLM accuracy by up to 18.0 points on CMM

Synopsis

Through path-intervention and representation analyses, this work identifies a question-relay mechanism in audio-visual large language models, where question-position hidden states carry both required-modality evidence and interfering-modality cues, and proposes SECRET, a training-free method that steers question states toward required-source evidence using source-conditioned positive and negative question representations, achieving higher accuracy than prior training-free methods on CMM and AVHBench across three AVLLMs, with gains of up to 18.0 and 7.1 percentage points over base models, and reducing distractor overlap in modality-specific captioning.

AI-generated editorial illustration: Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models

Interpretation

The paper identifies a question-relay mechanism: hidden states at question-token positions relay both required-modality evidence and interfering-modality cues into answer prediction, producing source-confused grounding hallucination. Prior work examined modality reliance at generation positions or mitigated the failure through inference-time corrective decoding and training-time alignment, leaving the internal cross-modal interactions insufficiently understood; this work moves the analysis to question states as an intermediate relay. Based on attention-path cutting and representation analyses on the Video-Driven Audio Hallucination and Audio-Driven Video Hallucination subsets of AVHBench: cutting required-modality pathways to question states reduces correct-answer support on source-faithful cases, cutting interfering-modality pathways to question states partially restores correct-answer support on source-confused cases, and question-position interventions yield greater correct-answer logit recovery than generation-position interventions; conclusions hold in direction across three- and five-layer windows.

The paper proposes SECRET, a training-free method that performs source-conditioned relay steering via contrasting question representations, suppressing cross-modal interference without altering audio-visual inputs. Unlike approaches that perturb or remove modality inputs, SECRET constructs positive and negative question representations through internal attention-path interventions, preserving the complete audio-visual context; unlike generation-position interventions, it targets question-token positions. The method first prompts the model with the textual question alone to predict the required modality, then cuts interfering-modality and required-modality pathways into question states to obtain positive and negative references, and updates original question states with a norm-matched, token-wise difference, applied only during prefill; ablations show that removing norm matching lowers accuracy on both models, and both the generation-position intervention variant and the modality-removal variant underperform SECRET.

Evaluated on CMM and AVHBench across three AVLLMs, SECRET consistently outperforms the compared training-free methods, improving overall accuracy over base models by up to 18.0 and 7.1 percentage points. The evaluation spans different model scales and both dense and mixture-of-experts architectures on two datasets, indicating that question-relay steering motivated by AVHBench analysis transfers across datasets. CMM uses the visual-dominance and audio-dominance subsets totaling 800 questions, and AVHBench's two subsets comprise 3,426 question-answer pairs; comparisons include training-free methods VCD, AVCD, and MAD. Gains vary with the required modality: VideoLLaMA2-AV-7B and Qwen2.5-Omni-7B improve more on audio-required tasks, while Qwen3-Omni-30B-A3B benefits more on video-required tasks.

On modality-specific captioning under mismatched audio-video inputs, SECRET achieves the lowest D-CIDEr and highest LLM-score while maintaining competitive T-CIDEr, indicating generalizability to open-ended generation. Main results center on discriminative question-answering benchmarks; this experiment extends source-grounding capability to open-ended description and measures both target agreement and distractor leakage. Following the ACPO evaluation setup, with 400 audio-swapped examples each for audio-target and video-target captioning; lower D-CIDEr indicates less overlap with distractor references, and higher LLM-score indicates better target fidelity and less distractor leakage; GPT-4.1 scoring agrees with human preference at 89% for audio-target and 87% for video-target captioning, versus human-human agreement of 93% and 90%.

Perspective

The results target audio-visual LLM inference settings where the instruction specifies the required modality, applicable to the source-confused grounding hallucination evaluations covered by CMM and AVHBench and to modality-specific captioning under mismatched audio-video inputs. The method is a training-free inference-time steering approach that requires no retraining, so it can be layered onto existing AVLLMs; the paper reports model-specific steering depths for VideoLLaMA2-AV-7B, Qwen2.5-Omni-7B, and Qwen3-Omni-30B-A3B, indicating that deployment requires selecting the intervention layer per model. For practitioners seeking to improve multimodal system reliability, this offers a path that retains the complete audio-visual context and corrects only at question states.

Required-modality identification relies on the model's own judgment from the textual question alone; the paper reports 99.85% accuracy for Qwen2.5-Omni-7B on this diagnostic, but randomly selects a modality when the prediction is ambiguous, and how this handling behaves on more complex questions remains an open question. Steering depth is set per model, and the paper observes that the best depth coincides with the lowest cosine similarity between positive and negative question representations; whether this association holds across more models and tasks awaits further verification. The mechanistic analysis focuses on pathway-level cross-modal information flow, and the paper notes that finer-grained analysis of individual attention-head roles has not yet been carried out. In addition, the modality-specific captioning evaluation uses 400 audio-swapped examples and GPT-4.1 scoring; although supported by human validation, the scope of samples and scoring criteria leaves room for extension.

Sources