acceptodds
Under review as a conference paper at ICLR 2027

Seen, Heard, or Assumed? Cross-Modal Hallucinations in Audio-Visual Language Models

Abstract

Despite the rapid progress of audio-visual large language models (AVLLMs), hallucination remains an important reliability challenge in joint audio-visual reasoning. Existing audio-visual hallucination benchmarks emphasize final-answer correctness or answer changes after modifying one modality, while it remains unclear whether these evaluation results reflect correct and stable judgments about what is heard and seen. To jointly assess final-answer correctness, modality-judgment correctness, and cross-modal stability, we introduce AV-TRACE, a benchmark with 7,344 test examples that combines explicit modality judgments with matched audio-visual comparisons for modality-specific hallucination detection and cross-modal stability testing. AV-TRACE reveals that correct final answers can mask modality-judgment failures, and that higher modality-level correctness can coexist with poorer cross-modal stability. Supervised fine-tuning improves answer and judgment accuracy, yet correct judgments can still flip when only the other modality changes. We therefore propose Factorial Quartet Fine-Tuning (FQFT), which augments GRPO with a joint answer-judgment reward and factorial regularization of judgment margins across four target-presence states. On Qwen2.5-Omni-7B, compared with supervised fine-tuning, FQFT increases the task joint accuracy of answers and modality judgments from 76.29% to 84.39%, reduces the hallucination rate from 20.49% to 10.32%, and reduces the flip rate from 4.08% to 1.95%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.