CIM Decomposes MLLM Uncertainty via Causal-Invariant Masking, with EED Proxy Matching Performance at Nearly 50% Speedup
Related research and updatesSynopsis
Addressing hallucinations in Multimodal Large Language Models (MLLMs), this work proposes Causal-Invariant Masking (CIM), which measures the semantic shift between original predictions and those conditioned on a causally-focused view to decompose uncertainty types, introduces Semantic Divergence as the core UQ metric with theoretical evidence that it converges to the variance of the model's sensitivity to non-causal correlations, and further proposes Expected Embedding Drift (EED), a fast geometric proxy metric; experiments report state-of-the-art performance on various benchmarks, with EED accelerating by nearly 50% at comparable performance.
Figure 1: Left: A case of blurred and background-mixup image pair. Middle: Sensitivity of Semantic Entropy to visual perturbations in MLLMs. Similar levels of error to the original image are binned into the same point. Right: Comparison of uncertainty increase under blurring and background mixup perturbations. While Semantic Entropy increases slightly across both, it fails to distinguish hallucination. In contrast, our proposed Semantic Divergence remains robust to visual blur but exhibits a distinct spike under hallucinative perturbations. The values are normalized by the standard deviation of each metric.
arXivInterpretation
Proposes the Causal-Invariant Masking (CIM) framework, which quantifies uncertainty by measuring the semantic shift between original predictions and predictions conditioned on a causally-focused view, thereby decomposing uncertainty types for more comprehensive UQ. Existing approaches are biased toward aleatoric uncertainty from data ambiguity and struggle to detect uncertainty caused by superficial associations when the query-relevant signal is weak; CIM instead targets epistemic uncertainty stemming from model limitations. The abstract-level text presents the framework design and motivation but provides no dataset, sample size, or ablation details.
Introduces Semantic Divergence as the core UQ metric and provides theoretical evidence that it converges to the variance of the model's sensitivity to non-causal correlations, establishing its ability to capture MLLM limitations. Links the uncertainty measure to a sensitivity-variance quantity under a causal view rather than relying only on empirical scoring. The abstract states theoretical evidence for convergence, but the visible text does not show proof details or assumptions.
Proposes Expected Embedding Drift (EED), a fast geometric proxy metric that estimates semantic shift directly within the hyperspherical embedding space to accelerate UQ in MLLMs. Replaces full semantic-shift computation with a geometric proxy, reducing computational cost while keeping performance comparable. The abstract reports nearly 50% acceleration with comparable performance but gives no benchmark names or numeric tables.
Experiments show the method achieves state-of-the-art performance on various benchmarks. Outperforms existing UQ approaches across multiple evaluation settings. An abstract-level overall claim; the benchmark list, comparison methods, and statistics are not provided in the visible text.
Perspective
The work targets multimodal large language model settings that require reliable deployment, especially cases where the query-relevant signal is weak and superficial associations readily induce hallucinations; its goal is to provide decomposable uncertainty signals so practitioners can distinguish data ambiguity from model-limitation uncertainty at inference time. As a geometric proxy, EED suits large-scale inference or online monitoring settings that need fast semantic-shift estimation in hyperspherical embedding space and are sensitive to computational cost.
The visible text is abstract-level and omits benchmark names, comparison methods, sample sizes, ablations, and the assumptions behind the theoretical proof, so the scope and robustness of the state-of-the-art conclusion cannot be judged; 'comparable performance' for EED lacks a quantitative definition, and the measurement conditions for the nearly 50% speedup (hardware, batch size, model scale) are unspecified; whether the causal assumptions underlying Semantic Divergence's convergence hold on real multimodal data remains an open question to watch.
