Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

RED, a training-free decoding method, raises accuracy by up to 7% on three audio-visual hallucination benchmarks at about 1.5x standard time to first token

The work introduces Relevant Evidence Decoding (RED), a training-free method that uses pointwise mutual information to decompose the joint audio-visual prediction into audio, video, and residual interaction components, then runs a question-only inference pass to determine the required evidence type and augments the original prediction with the corresponding contribution; across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding.