Perception or Prior? Motion-Blind Controls for Video Hallucination Mitigators
Abstract
Hallucination mitigators for video LLMs are judged by benchmark deltas, yet many benchmark questions are answerable without temporal information, and preference optimization can raise scores by shifting language priors alone. How much of any reported gain requires motion remains unmeasured. We provide the measurement. Five DPO arms share one base model, one public preference dataset, and one training recipe, differing only in the visual signal seen during training: the original video, uniformly permuted frames, black frames, no visual tokens, or order-preserving jitter. Arm equivalence is established by recipe digests and per-run consumption hashes. A preregistered linear contrast adjudicates the threshold claim without division; recovered-gain fractions are reported behind a denominator gate with Fieller intervals, conditioned on a motion-necessary stratum identified by five held-out model families. We report a rank order of bounds, not an additive decomposition. Two evaluation-side findings stand on their own. Freezing every frame to a repeated median costs Qwen2.5-VL-7B 13.3 VideoHallucer pair-accuracy points; black frames cost 43.1—the blind control overstates the video's contribution roughly threefold. Separately, switching answer extraction from generative parsing to per-option log-likelihood moves Video-LLaVA-7B 15 points (9.5 to 24.5), bracketing the published 17.8 and exceeding the margin between most mitigators and their baselines. Neither benchmark specifies which extraction rule was used. Checkpoints, the motion-necessary stratum, and all intervention code are released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.