acceptodds
Under review as a conference paper at ICLR 2027

Vid-Evi-QA: A Controlled Benchmark for Evidence-Adaptive Video Reasoning

Abstract

Existing video question answering (VideoQA) benchmarks assess whether models answer or refuse correctly, but not whether decisions adapt to changes in relevant visual evidence. We introduce Vid-Evi-QA, a controlled evidence-intervention benchmark spanning Sufficient, Insufficient, Partial, and Conflicting evidence, with 24,357 condition-level instances from 5,422 questions and 1,739 videos. It intervenes on temporal evidence while holding questions fixed and separately constructs cases with multiple visually supported options. Same-question pairs and matched non-evidence controls distinguish answering ability, refusal tendency, and evidence sensitivity. Evaluations of ten open-weight video-language models across all four evidence conditions reveal a clear separation between these capabilities: the strongest answerer attains 70.53% Sufficient accuracy but only 1.87% strict Evidence Withdrawal Rate (EWR; correct-reason withdrawal conditional on initial correctness) and 1.32% joint Evidence-Contingent Success (ECS); the highest Evidence-Specific Success in matched-control experiments is just 12.68%. Unstable decisions under graded evidence and failures to recognize and appropriately reject multiple supported options reveal further limitations. We also introduce two training-free inference methods that explicitly assess option-level evidence support. Across all 9,209 non-Partial instances, two-pass visual verification improves the reason-correct paired success of the backbone from 1.45% to 18.01% and conflict detection from 5.44% to 50.95%. These results demonstrate the potential of explicit support assessment to improve evidence-aware perception and decision making, and show that video reasoning evaluation must go beyond static accuracy and refusal rates to directly test whether decisions adapt appropriately to changes in supporting evidence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.