RED, a training-free decoding method, raises accuracy by up to 7% on three audio-visual hallucination benchmarks at about 1.5x standard time to first token
Related research and updatesSynopsis
The work introduces Relevant Evidence Decoding (RED), a training-free method that uses pointwise mutual information to decompose the joint audio-visual prediction into audio, video, and residual interaction components, then runs a question-only inference pass to determine the required evidence type and augments the original prediction with the corresponding contribution; across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding.
Figure 1: Motivation of this study. (a) A correct prediction does not necessarily indicate successful hallucination mitigation. (b) When question-required evidence changes, the prediction should change accordingly. (c) When non-required evidence changes, it should remain unchanged. The base model violates both behaviors, whereas RED follows the evidence required by the question.
arXivInterpretation
The paper observes that joint audio-visual inference can weaken a prediction even when a single modality already yields the correct answer, for example correctly predicting violin from audio alone while confidence in violin drops once a video showing a guitar is added. Prior contrastive decoding targeted vision-language models, and extending it directly to AV-LLMs overlooks that different questions require different perceptual evidence: audio, video, or their interaction. The observation is presented in the abstract as a motivating example; no statistics or sample sizes are reported.
RED uses pointwise mutual information to quantify the predictive support that audio and video provide beyond the question alone, decomposing their joint contribution into audio, video, and residual interaction components. Unlike a direct contrastive-decoding extension, RED explicitly separates evidence sources so that strengthening acts on the question-relevant modality component. The method is described as training-free, with the abstract giving the PMI decomposition and question-only pass as the two-step procedure; implementation details and ablations are not provided.
RED first performs a question-only inference pass to determine the required evidence type, then augments the original audio-visual prediction with the corresponding PMI contribution. This selection mechanism makes 'which evidence does the question need' a precondition for decoding rather than weighting all modalities uniformly. The abstract states the order of steps but does not report the accuracy of this determination or its failure cases.
Across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding. The results span multiple benchmarks and models while also reporting the inference-latency cost, allowing accuracy and efficiency to be weighed together. The abstract reports benchmark names, the number of models, the improvement magnitudes, and relative latency, but no per-model breakdown or significance testing.
Perspective
The work targets cross-modal hallucination mitigation in audio-visual large language models, applying to audio-visual question answering where evidence must be drawn from audio, video, or their interaction, and it plugs into existing AV-LLMs at the decoding stage without training. Beneficiaries include researchers studying multimodal hallucination and decoding strategies, and practitioners who want to reduce audio-visual hallucination without large-scale retraining. Results are scoped to three audio-visual hallucination benchmarks (CMM, AVHBench, SVHalluc) and three AV-LLMs, with latency framed as an average time to first token of about 1.5x standard decoding.
A careful reader would still watch how reliably the question-only pass identifies the evidence type and on which questions it fails; when the residual interaction component dominates the PMI decomposition; how the 1.5x time to first token plays out in long-output or real-time settings; and whether the improvement magnitudes are consistent across the different AV-LLMs. The available text is abstract-level, lacking figures, per-model results, and ablation details, so these questions cannot be settled from the present material and remain open for a full reading.
