Seen, Heard, or Assumed? Cross-Modal Hallucinations in Audio-Visual Language Models
Abstract
Despite the rapid progress of audio-visual large language models (AVLLMs), hallucination remains an important reliability challenge in joint audio-visual reasoning. Existing audio-visual hallucination benchmarks emphasize final-answer correctness or answer changes after modifying one modality, while it remains unclear whether these evaluation results reflect correct and stable judgments about what is heard and seen. To jointly assess final-answer correctness, modality-judgment correctness, and cross-modal stability, we introduce AV-TRACE, a benchmark with 7,344 test examples that combines explicit modality judgments with matched audio-visual comparisons for modality-specific hallucination detection and cross-modal stability testing. AV-TRACE reveals that correct final answers can mask modality-judgment failures, and that higher modality-level correctness can coexist with poorer cross-modal stability. Supervised fine-tuning improves answer and judgment accuracy, yet correct judgments can still flip when only the other modality changes. We therefore propose Factorial Quartet Fine-Tuning (FQFT), which augments GRPO with a joint answer-judgment reward and factorial regularization of judgment margins across four target-presence states. On Qwen2.5-Omni-7B, compared with supervised fine-tuning, FQFT increases the task joint accuracy of answers and modality judgments from 76.29% to 84.39%, reduces the hallucination rate from 20.49% to 10.32%, and reduces the flip rate from 4.08% to 1.95%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.