Eyes Over Ears: Mitigating Visually Induced Acoustic Hallucinations in Omni-Modal LLMs
Abstract
Omni-modal LLMs jointly process language, vision, and audio. In practice, audio is incorporated only after a vision-language backbone has been trained. We find that established visual priors in such backbones can override acoustic evidence on audio-related tasks, giving rise to *Visually Induced Acoustic Hallucinations*. To diagnose and mitigate these hallucinations, we develop a diagnostic benchmark and synthesize training data targeting these failure modes. Directly fine-tuning existing open-source omni-modal LLMs on these data yields limited changes to their learned modality-dependence patterns, motivating intervention earlier in training. We therefore introduce Audite-Omni-4B, a compact model trained through progressive audio alignment, Mid-Training, and SFT. Introducing targeted supervision from Mid-Training onward proves more effective, shifting multimodal attention toward audio tokens during audio-related reasoning and substantially improving performance on our diagnostic benchmark. The improvements also extend to open-source omni QA and captioning tasks, indicating generalization beyond the targeted failure mode. We further train task-specialized GRPO experts and consolidate their respective strengths into a single model through MOPD. Audite-Omni-4B scores 89.9 on our diagnostic suite, compared with 72.7 for the strongest closed-source baseline Qwen3.8-Omni-Flash. It also surpasses Qwen3-Omni-30B-A3B on 10 of 11 public omni QA and captioning benchmarks, establishing a new open-source SOTA.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.