Beyond Rewatching: Learning Evidence-Driven Complex Video Reasoning
Abstract
Complex video reasoning requires integrating visual evidence across entities, events, and time. Although recent models interleave textual reasoning with visual inspection, tool use alone does not ensure evidence-driven reasoning: models may inspect unnecessarily, fail to adapt unsuccessful searches, or retain interpretations contradicted by new observations. We address these limitations through three complementary training improvements. First, we propose Evidence-Conditioned Synthesis (ECS), which constructs verified reasoning continuations from intermediate states in executed video interactions, providing supervision for evidence assessment, selective inspection, recovery, and belief updating. Second, we introduce EviVideo-QA-20K, a dataset of 20K perception-centric questions that require composing visual evidence through conjunctive and sequential reasoning. Third, we design an Evidence Reward that complements answer correctness with an assessment of evidence use and an explicit inspection cost. Together, these components train EviVideo through supervised fine-tuning followed by reinforcement learning. Our 4B and 8B models improve over their Qwen3-VL backbones by average absolute accuracy gains of 7.6% and 7.9% across MINERVA, Minerva-Ego, and Video-Holmes, and by 4.0% and 3.6% across Video-MME and Video-MME-v2. Behavioral evaluations further indicate improved evidence-responsive inspection and revision. These results highlight the importance of teaching models not merely to invoke visual tools, but to let evidence guide reasoning and interaction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.