Reliable Reasoning for Video Understanding
Abstract
Multimodal chain-of-thought (CoT) has emerged as a promising paradigm for improving the reasoning capabilities of multimodal large language models (MLLMs): It typically extracts question-relevant cues from a static image and reasons over them to derive an answer. The success of this paradigm partly rests on the assumption that the static image contains sufficient evidence for the question. Videos, however, challenge this assumption because they are temporally structured sequences of frames in which relevant evidence can be distributed unpredictably over time. This introduces two sources of unreliability. First, a single sampled video observation may omit crucial evidence, causing the model to reason from an under-informed view; we refer to this as unreliability arising from under-informed reasoning. Second, the model's understanding and answer may change substantially as richer observations become available; we refer to this as unreliability arising from cross-observation inconsistency. To address both challenges, we introduce Reliable Reasoning with Videos (), which assesses answer reliability at two levels. Within a single observation, RRV uses video information extraction and answer reasoning spans when both are available, falling back to whole-response uncertainty otherwise, and aggregates the reliability-weighted trajectories into an answer distribution. Across observations, RRV checks answer consistency along a nested sequence of increasingly informative frame sets until two adjacent observations propose the same answer. By estimating the reliability of CoT within each observation and filtering answers based on consistency across observations, RRV suppresses the influence of unreliable CoT reasoning on final predictions. Experiments on three video MLLMs and eight video QA benchmarks show that RRV improves accuracy in most evaluated settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.