Tri-PvP's 8,000 tri-modal conflict samples show omni-modal LLMs favor images, and favor perceptual vision but propositional audio
Synopsis
The authors built Tri-PvP, an 8,000-sample tri-modal conflict benchmark spanning animal, emotion, environment, and music domains in which image and audio each appear in perceptual form (real photographs or recordings) or propositional form (text-card images or TTS speech) while text is always propositional, evaluated five omni-modal LLMs across four evidence-type conditions, and found that image bias dominates in 18 of 20 model-by-condition bars and that a systematic asymmetry exists: models favor perceptual evidence in vision but propositional evidence in audio, with this bias linearly decodable from early hidden layers and only partially mitigated by contrastive decoding, which introduces a residual text bias.
Interpretation
The paper identifies a structural confound in existing tri-modal bias benchmarks: perceptual evidence (a photograph or recording of a dog) and propositional evidence (the declarative claim "this is a dog") are conflated within a single modality, while vision is overwhelmingly natural images, audio mixes real recordings with synthesized speech, and text is propositional by construction, so any measured modality bias is entangled with evidence-form bias and cannot be cleanly attributed. Relative to prior work, the paper formalizes the epistemological distinction between perceptual and propositional evidence as an experimentally manipulable axis and designs a benchmark around it, letting image and audio each take perceptual or propositional form so that modality bias and evidence-form bias can be separated. The claim rests on the benchmark design itself: four evidence-type conditions (PercI-PercA, PercI-PropA, PropI-PercA, PropI-PropA) are constructed as matched counterfactuals, reusing the same asset whenever a channel takes the same evidence form, with the number of triples balanced across cross-modal label combinations.
Across five omni-modal LLMs, image bias dominates in most models and conditions, audio bias is consistently the least pronounced single-modality bias, and switching the visual or audio channel from perceptual to propositional only reduces, never reverses, this preference. The paper quantifies direction and magnitude: BIAS_IMAGE is the dominant label in 18 of 20 model-by-evidence-type bars and frequently exceeds 60%, while BIAS_AUDIO is the smallest of the three single-modality biases in 18 of 20 bars, typically below 10%. Evidence comes from evaluating Qwen2.5-Omni-7B, MiniCPM-o 4.5, Qwen3-Omni-30B-A3B-Thinking, Gemma 4 E4B, and Gemini 3 Flash, with free-form responses classified by an LLM judge into eight mutually exclusive labels and manual verification by the authors on 800 samples (10%) reaching 97.1% agreement.
The paper reveals a systematic asymmetry in evidence-form bias: models show stronger bias toward perceptual evidence in vision but propositional evidence in audio; for example Qwen2.5's BIAS_IMAGE is 49.7% under PercI-PropA but only 12.7% under PropI-PropA, while Gemma4's audio bias rises from 0.9% under PercI-PercA to 25.4% under PercI-PropA. This asymmetry was not observable in prior benchmarks because it requires independently varying the evidence form of each channel; the paper links it to divergent pretraining objectives, with vision encoders generally trained to capture perceptual image content and audio encoders frequently initialized from ASR-style objectives emphasizing linguistic content recovery from speech. Evidence is the cross-model, cross-condition bias distribution reported with 95% confidence intervals; the authors present the encoder-objective explanation as a hypothesis and point to a need for stronger perceptual audio understanding in omni-modal models.
Layer-wise linear probing shows modality-bias information is already linearly decodable from intermediate hidden states before any token is generated, with image bias most decodable and audio least; contrastive decoding as an inference-time diagnostic reduces image bias and raises unbiased responses but introduces a residual text bias, while overall OmniBench accuracy moves only from 38.4% to 37.4%. This moves bias from an output-level phenomenon to a representation-level one and indicates surface-level interventions are insufficient; contrastive decoding cuts image bias from 74.9% to 38.8% and raises NO_BIAS from 7.2% to 38.6% under PropI-PropA, yet text bias rises in all four conditions by 2.2% to 13.5%. Probing trains per-layer logistic regressions on Gemma4 and Qwen2.5 and reports balanced accuracy, with Qwen2.5's image probe reaching 80% peak accuracy under both PercI conditions; the contrastive decoding experiment is confined to Gemma4 with fixed alpha and adaptive plausibility threshold, and the OmniBench evaluation uses 1,142 tri-modal multiple-choice items.
Perspective
This work targets the everyday perceptual scale, the scenarios where humans typically turn to direct perception, covering animal, emotion, environment, and music domains, which matches the deployment context of current omni-modal models as personal assistants; questions are deliberately bias-neutral so as to elicit unprompted modality preference. It is meant for researchers and developers who need to evaluate or diagnose which stream an omni-modal model trusts when image, audio, and text conflict, and for follow-up work that reuses its controlled evidence-form design to build new benchmarks. Contrastive decoding is positioned as a diagnostic inference-time intervention requiring no parameter updates.
The boundaries the paper itself notes are worth watching: the benchmark covers only the everyday scale, so how models behave under multi-modal conflicts at broader epistemic scales such as complex diagrams or scientific data remains unclear; questions are deliberately bias-neutral, and how preference shifts under prompts that request enumeration of all sources, demand a single answer, or supply source-reliability cues is left to future work. The contrastive decoding experiments are confined to a single model with fixed alpha and adaptive plausibility threshold, without a systematic sensitivity analysis or evaluation across a broader set of models, and they require an additional forward pass, making them computationally more expensive. In addition, the composition of NO_BIAS responses varies substantially across models, ranging from genuine conflict acknowledgment to passive enumeration of modality contents, so the claim that models are less prone to prioritizing a single input stream under propositional conditions needs to be read per model. This summary is based on the paper text and its appendices and does not include the image content of Figures 2 through 6.
