Learning to Reason Efficiently in Videos: Reinforcement Learning with Short- and Long-Term Rewards
Abstract
Recent advancements in multimodal reasoning models have shown promising results, but these models often inherit a language-centric reasoning bias and fall short in true multimodal understanding. In long video understanding, a core challenge in multimodal reasoning, such models tend to generate lengthy and reflective chain-of-thought (CoT) outputs that do not meaningfully contribute to video comprehension. Moreover, the excessive token generation during reasoning leads to inefficiencies in further training. To address these issues, we propose a simple post-training framework for building efficient video reasoning models. Specifically, to mitigate the text-biased reasoning patterns inherited from language models, we first distill an Instruction multimodal model using the video reasoning trajectories of a native multimodal reasoning model, allowing it to learn dynamic video reasoning. During the reinforcement learning stage, to improve data efficiency, we leverage the instruction-tuned model to filter long-form video question-answering data, retaining only a compact set of high-quality training samples. For reward design, we introduce a hybrid mechanism that incorporates both short-term and long-term rewards, optimized via group relative policy optimization (GRPO). Extensive experiments on diverse video benchmarks show that our framework transforms Qwen3-VL-Instruct into a more accurate and efficient video reasoning model, substantially outperforming the open-source Qwen3-VL-Thinking baseline. Consistent gains on Qwen2.5-VL and Qwen3.5 further demonstrate its generalizability. These results highlight the importance of dynamic reasoning and efficient reward design for advancing multimodal video reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.