CF-GRPO: Counterfactual Group Relative Policy Optimization for Video Hallucination Suppression
Abstract
Video language models have made substantial progress in video understanding, yet they remain prone to hallucination, producing answers that are inconsistent with the video content. Reinforcement fine-tuning with verifiable outcome rewards has been explored to mitigate such hallucinations. However, these rewards cannot distinguish visually grounded correct answers from those produced by the language prior. We show that this source-blind credit assignment is inherent in outcome-reward GRPO: the expected update is determined by the probability of correctness, regardless of whether that correctness depends on visual evidence. Our analysis further motivates interventions at three complementary levels of the GRPO update dynamics: the direction of pair-level updates, the magnitude of sample-level reinforcement, and the allocation of credit within each output group. Accordingly, we propose CF-GRPO, with three mechanisms targeting these respective levels. Specifically, 1)At the pair level, counterfactual pair optimization jointly trains opposite-label prompts sharing the same video, cancelling the common answer-shift update of prompts that the policy answers alike while preserving a positive discriminative update. 2)At the sample level, blind-failure advantage shaping scales the GRPO advantage by the model's failure probability without the video, reducing reinforcement for samples already solvable from the language prior. 3)At the within-group level, an evidence-consistency reward distinguishes outputs with identical correctness by whether their reasoning states answer-consistent evidence, restoring learning signal when outcome-reward GRPO provides none. Experiments on three video hallucination benchmarks show that CF-GRPO consistently improves over GRPO and strong baselines, while preserving the model’s answer distribution and general video understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.