acceptodds
Under review as a conference paper at ICLR 2027

CAD: Conflict-Aware Decoding to Mitigate Cross-Modal Hallucinations in Omnimodal Large Language Models

Abstract

Omnimodal large language models (Omni-LLMs) integrate audio, video, and text for perception and reasoning but remain vulnerable to cross-modal hallucinations, where one modality improperly influences predictions about another. By avoiding the additional parameter updates required by training-based methods, training-free decoding offers an appealing approach to mitigating these hallucinations. Existing training-free methods modulate modality contributions during decoding through perturbation or relevance weighting, but do not explicitly assess fusion effects on joint predictions, potentially retaining harmful interference or suppressing useful complementarity. To this end, we propose Conflict-Aware Decoding (CAD), a training-free, two-stage framework. First, to assess the potential effects of cross-modal fusion, Potential Conflict Magnitude Estimation (PCME) quantifies audio-video disagreement and the joint prediction's deviation from a relevance-weighted unimodal reference. Second, to assess evidential support for intervention, Conflict Actionability Assessment (CAA) builds on Dempster-Shafer reliability discounting to evaluate task-space answer relations using query relevance and answer decisiveness. When warranted, CAD shifts decoding weight from the joint branch to the unimodal branches. Extensive experiments demonstrate that CAD outperforms competitive training-free methods. It achieves substantial gains on cross-modal hallucination benchmarks CMM and AVHBench while also improving general audio-visual question answering on WorldSense and VideoMME.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.