TAG-GRPO: Mitigating Cross-Modal Hallucinations via Task-Aware Grouped GRPO
Abstract
Large omni-modal models jointly process audio, video, and language, yet they often produce correct answers through incorrect modality reasoning—relying on spurious correlations rather than task-relevant evidence, which results in cross-modal hallucinations in audio-visual understanding. Existing modality-aware preference optimization methods mitigate this issue via offline perturbation pairs; however, their static supervision struggles to capture task-dependent modality roles. To address this, we propose , an online policy optimization framework that extends GRPO from standard response-only rollouts to task-aware, modality-conditioned rollouts. Specifically, TAG-GRPO generates task-aware modality conditions through diffusion-based corruption and routes them into Intact, Robustness, and Sensitivity groups according to the task type. Group-wise advantage normalization then independently optimizes correctness under full evidence, invariance to irrelevant modality perturbations, and sensitivity to the disruption of task-relevant evidence. This explicit behavioral separation encourages the model to ground its reasoning in proper modality evidence, thereby suppressing cross-modal hallucinations beyond what answer-level correctness alone can achieve. Extensive experiments on AVHBench and CMM demonstrate that TAG-GRPO improves audio-visual grounding and reduces cross-modal hallucinations, highlighting the critical importance of optimizing modality-faithful behavior beyond answer correctness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.