GETR: Guided Exploration Trajectories for Long-Video Reasoning
Abstract
Long-video reasoning requires finding and combining evidence that may appear throughout an hour-long video. An answer-only reward does not tell the policy whether it inspected the supporting evidence, because the same answer may come from partial or irrelevant observations, or from a guess. Giving the model a tool to look closer does not help by itself, because an untrained model does not know where to look: Qwen3-VL-8B with a crop tool is less accurate than with uniform frames. We propose Guided Exploration Trajectories for Reasoning (GETR), which rewards a long-video policy for looking at the correct evidence. The policy starts from hints, candidate segments retrieved offline from text summaries of the video, and is rewarded for where its first crop falls and for whether its final crop covers the labeled evidence. An IoGT coverage reward measures whether its final crop contains the labeled evidence, and we report how often any crop reaches that evidence. We train the model with GRPO on 1,668 question–answer–evidence triplets, without SFT or labeled tool calls, whereas LongVT and EVA fine-tune on 263K and 21K supervised samples, respectively, including distilled tool-call trajectories. Across four long-video benchmarks (Video-MME-Long, LVBench, HERBench-Lite, and CG-Bench), GETR improves the Qwen3-VL-8B base model by 3.6 points on average, and by 4.4 on which the same model given the crop tool without training instead loses 5.8 points. The gains are largest where the evidence is short relative to the video (CG-Bench +6.3, LVBench +5.4).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.