Public articles linked to the same research event.
arXiv The work introduces Relevant Evidence Decoding (RED), a training-free method that uses pointwise mutual information to decompose the joint audio-visual prediction into audio, video, and residual interaction components, then runs a question-only inference pass to determine the required evidence type and augments the original prediction with the corresponding contribution; across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding.
The work introduces Relevant Evidence Decoding (RED), a training-free method that uses pointwise mutual information to decompose the joint audio-visual prediction into audio, video, and residual interaction components, then runs a question-only inference pass to determine the required evidence type and augments the original prediction with the corresponding contribution; across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding.
The work introduces Relevant Evidence Decoding (RED), a training-free method that uses pointwise mutual information to decompose the joint audio-visual prediction into audio, video, and residual interaction components, then runs a question-only inference pass to determine the required evidence type and augments the original prediction with the corresponding contribution; across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding.
The work introduces Relevant Evidence Decoding (RED), a training-free method that uses pointwise mutual information to decompose the joint audio-visual prediction into audio, video, and residual interaction components, then runs a question-only inference pass to determine the required evidence type and augments the original prediction with the corresponding contribution; across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding.