EGAA: Evidence-Grounded Answer Alignment for Mitigating Cross-Modal Hallucination in Omni-Modal LLMs
Abstract
Cross-modal hallucination remains a major challenge for omni-modal large language models. Existing benchmarks such as AVHBench and CMM measure these failures at the task level, but their standard scores provide limited information about the specific error within a generated response. We therefore examine generated answers at two complementary levels. Cross-Modal Source Misattribution (CSM), Absence Overgeneralization (AO), and Unsupported Certainty (UC) are identified at the statement level, while Wrong-Modality Answer (WMA) is assessed at the response level. To support both analysis and training, we augment AVHBench and CMM separately with visual and audio evidence annotations, supported statements, grounded answers, and error-targeted negative answers for these four error types. Based on these annotations, we propose Evidence-Grounded Answer Alignment (EGAA), a parameter-efficient post-training method that jointly learns from modality-specific evidence, supported statements, grounded answers, benchmark-format targets, and constructed preference pairs. We evaluate EGAA through bidirectional transfer between AVHBench and CMM. With Qwen2.5-Omni-7B, training on AVHBench improves CMM accuracy from 81.1% to 88.8%, while training on CMM improves AVHBench overall accuracy from 76.1% to 82.2%. These results show that explicitly connecting generated statements to multimodal evidence can reduce cross-modal hallucination across benchmarks. The code and configuration files used to reproduce our results are available at https://anonymous.4open.science/r/EGAA-DF71/, and the corresponding evidence annotations and dataset records are available at https://osf.io/j7x89/overview?view_only=2ff822a2fee8471f9b6384c25d00ecf6.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.