acceptodds
Under review as a conference paper at ICLR 2027

TAG-GRPO: Mitigating Cross-Modal Hallucinations via Task-Aware Grouped GRPO

Abstract

Large omni-modal models jointly process audio, video, and language, yet they often produce correct answers through incorrect modality reasoning—relying on spurious correlations rather than task-relevant evidence, which results in cross-modal hallucinations in audio-visual understanding. Existing modality-aware preference optimization methods mitigate this issue via offline perturbation pairs; however, their static supervision struggles to capture task-dependent modality roles. To address this, we propose , an online policy optimization framework that extends GRPO from standard response-only rollouts to task-aware, modality-conditioned rollouts. Specifically, TAG-GRPO generates task-aware modality conditions through diffusion-based corruption and routes them into Intact, Robustness, and Sensitivity groups according to the task type. Group-wise advantage normalization then independently optimizes correctness under full evidence, invariance to irrelevant modality perturbations, and sensitivity to the disruption of task-relevant evidence. This explicit behavioral separation encourages the model to ground its reasoning in proper modality evidence, thereby suppressing cross-modal hallucinations beyond what answer-level correctness alone can achieve. Extensive experiments on AVHBench and CMM demonstrate that TAG-GRPO improves audio-visual grounding and reduces cross-modal hallucinations, highlighting the critical importance of optimizing modality-faithful behavior beyond answer correctness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.