Test-Time Paired Option-Wise Verification to Reduce Prior-Following Errors in VideoLLMs
Abstract
Video large language models (VideoLLMs) lose accuracy when only a few frames are available. We found that this degradation stems not only from missing visual evidence but also from prior-following errors: driven by its language prior, the model gives the same wrong answer with or without the video. Replacing the original video with an unrelated one often preserves the preferred option, indicating weak dependence of the decision on the question-specific video. However, reversing the temporal order changes the relative support for the correct and no-video preferred options, showing that sampled frames retain discriminative evidence. This mismatch between retained video evidence and prior-aligned decisions motivates Paired Option-Wise Verification (POV), a training-free correction method that independently verifies each answer option with and without the video. To remove the language prior, POV subtracts a scaled no-video score from each option's video-conditioned score. The scale is estimated from unlabeled target-task data and is reduced only when the video evidence is consistent with the prior, that is, when the verifier supports the prior-preferred option. Across several frozen backbones, POV outperforms competing training-free baselines on temporal multiple-choice and binary QA, and even surpasses dense-frame VideoLLMs while using the sparse-frame budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.