acceptodds
Under review as a conference paper at ICLR 2027

Learning to Reason Efficiently in Videos: Reinforcement Learning with Short- and Long-Term Rewards

Abstract

Recent advancements in multimodal reasoning models have shown promising results, but these models often inherit a language-centric reasoning bias and fall short in true multimodal understanding. In long video understanding, a core challenge in multimodal reasoning, such models tend to generate lengthy and reflective chain-of-thought (CoT) outputs that do not meaningfully contribute to video comprehension. Moreover, the excessive token generation during reasoning leads to inefficiencies in further training. To address these issues, we propose a simple post-training framework for building efficient video reasoning models. Specifically, to mitigate the text-biased reasoning patterns inherited from language models, we first distill an Instruction multimodal model using the video reasoning trajectories of a native multimodal reasoning model, allowing it to learn dynamic video reasoning. During the reinforcement learning stage, to improve data efficiency, we leverage the instruction-tuned model to filter long-form video question-answering data, retaining only a compact set of high-quality training samples. For reward design, we introduce a hybrid mechanism that incorporates both short-term and long-term rewards, optimized via group relative policy optimization (GRPO). Extensive experiments on diverse video benchmarks show that our framework transforms Qwen3-VL-Instruct into a more accurate and efficient video reasoning model, substantially outperforming the open-source Qwen3-VL-Thinking baseline. Consistent gains on Qwen2.5-VL and Qwen3.5 further demonstrate its generalizability. These results highlight the importance of dynamic reasoning and efficient reward design for advancing multimodal video reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.