OmniConfess uses token-level channel-wise evidence attribution to mitigate omni-modal hallucination across text, image, audio, and video settings
Related research and updatesSynopsis
The work introduces OmniConfess, a training-free method that fixes a candidate response and re-scores it at token resolution under controlled channel-wise evidence interventions, yielding a structured token-by-channel confession that reveals the response's evidential dependence; it uses this confession to preserve grounded content and correct commitments driven by irrelevant or contradictory evidence, and it constructs OmniHalluBench, a 3,540-example benchmark built from six datasets spanning text, image, audio, and video settings and both judgment and free-form generation, with experiments showing hallucination mitigation across heterogeneous modality and task settings.
Figure 1: Examples of OmniLLM hallucination.
arXivInterpretation
It proposes OmniConfess, a training-free method for mitigating omni-modal hallucinations that fixes a candidate response and re-scores it at token resolution under controlled channel-wise evidence interventions, producing a structured token-by-channel confession that reveals the response's evidential dependence. Existing inference-time methods can reduce hallucinations but rarely reveal which evidence sustains a generated commitment; this work makes evidential dependence explicit as a readable per-token, per-channel structure. A method design described at the abstract level, emphasizing training-free inference-time intervention; implementation details, intervention construction, and scoring formulas are not expanded in the provided text.
Using this confession, OmniConfess preserves grounded content and corrects commitments driven by irrelevant or contradictory evidence. It turns attribution from a diagnostic signal into an actionable basis for correction, separating content that is grounded from content driven by the wrong evidence. A stated method objective from the abstract; no quantitative before-and-after figures are provided.
It constructs OmniHalluBench, a 3,540-example benchmark built from six datasets spanning text, image, audio, and video settings and both judgment and free-form generation. It provides a unified test collection for omni-modal hallucination evaluation across modalities and task forms, addressing the tendency of prior evaluation to focus on a single modality or task form. The abstract explicitly states the 3,540 example count and six datasets, and describes the covered modalities and task types.
Experiments show that OmniConfess mitigates hallucinations across heterogeneous modality and task settings, and the code and benchmark are publicly available. It validates applicability across multiple modality and task combinations and releases code and benchmark to support reproduction and subsequent comparison. An abstract-level overall experimental conclusion; no specific metric values, baseline names, or effect sizes are given.
Perspective
The work targets hallucination at inference time in omni-modal large language models, applicable to mixed text, image, audio, and video inputs and to both judgment and free-form generation task settings; being training-free, it can be attached to an inference pipeline without retraining the model. Its attribution output serves users who need to know which channel evidence a response depends on, such as evaluation and debugging scenarios that examine the basis of model statements. The released code and OmniHalluBench benchmark enable subsequent comparisons on the same test collection.
The provided text is at the abstract level and does not list specific evaluation metrics, baseline systems, or per-item results, so the magnitude of mitigation and relative strength across modalities cannot be judged. The exact construction of channel-wise evidence interventions, the computation details of token-resolution re-scoring, and how attribution is converted into preservation or correction decisions all require the main text. OmniHalluBench is built from six existing datasets, and its sample composition and task distribution are not described in the abstract, so consistency of conclusions across datasets awaits main-text data. In addition, the abstract does not state the model scales or modality combinations under which the method fails, and this scope question can serve as a point for later observation.
