GroundBoth: Support-Counter Grounding for Video Question Answering
Abstract
Grounded video question answering (VideoQA) asks a model to support its answer with localized temporal evidence. However, a correct answer with accurately grounded evidence does not by itself show that the model understood the video. When a wrong option describes an event that also occurs in the video and resembles the correct one, a model can choose correctly without ever telling the two apart. We extend grounded video question answering to support–counter grounded VideoQA: a model grounds the selected answer and, for each rejected option, grounds the competing event it refers to and explains why that event does not answer the question, making each rejection inspectable rather than taken on trust. To study this task, we build SC-Ground by refining NExT-GQA, ReXTime, and CG-Bench so that every distractor is anchored to a real event elsewhere in the video, with support and counter annotations. Used as a diagnostic, it reveals that even strong proprietary models ground the selected answer well yet largely miss the events behind the wrong options; grounding the answer alone is far from enough. We therefore curate a support–counter training set and introduce GroundBoth, a 2B model that emits the answer, its support, and per-option counter evidence in a single pass, trained by role-wise on-policy distillation (Role-wise OPD) and support–counter GRPO (SC-GRPO) that rewards both evidence roles. Despite its size, GroundBoth rivals 7B models on the original benchmarks, raising grounded answering by 18.0, 15.5, and 2.7 points over its 2B backbone, and leads all evaluated open-source single-call models on SC-Ground, lifting joint grounding nearly eightfold on NExT-GQA-SC (1.48 to 11.67).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.