Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
Abstract
Recent advances in multimodal reasoning, particularly Thinking with Images, have achieved remarkable progress by empowering Multimodal Large Language Models (MLLMs) to actively inspect and reason over visual context. Extending this visual reasoning paradigm to dynamic videos requires models to actively localize and synthesize visual evidence across the temporal dimension. However, enabling MLLMs to think with videos faces two fundamental challenges: (i) from the data perspective, existing training data lacks tasks that guide models to actively seek and verify evidence across temporal intervals; and (ii) from the training supervision perspective, current methods lack verified temporal process supervision, leading to unfaithful reasoning and superficial shortcuts. To address these challenges, we propose Video-Thinker to unify data construction and training supervision, enabling intrinsic temporal reasoning without relying on external tools. From the data perspective, to formally capture both temporal and semantic dependencies across events, we construct a video reasoning graph to model their temporal ordering and semantic relations. Leveraging this graph, we synthesize reasoning tasks with guaranteed evidence necessity, ensuring that correct answers strictly require observing and integrating all designated temporal clues. From the training supervision perspective, following a cold-start Supervised Fine-Tuning (SFT) phase, we introduce an evidence-aware Reinforcement Learning (RL) method that provides verified temporal process supervision to ensure faithful grounding throughout reasoning. Extensive experiments across six challenging video reasoning benchmarks demonstrate that Video-Thinker achieves an average improvement of +6.86% over the base model with remarkable data efficiency, requiring only 20K training instances.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.