acceptodds
Under review as a conference paper at ICLR 2027

Route2Learn: Native Tool Routing with a Single MLLM for Long Video Understanding

Abstract

Long-form video understanding requires effective evidence acquisition strategies, as relevant information is often sparsely distributed across extended temporal spans. Different queries can benefit from different strategies: generation-based strategy excels at reasoning over global context, while retrieval-based strategy is better suited for locating specific events. However, existing methods either rely on a single strategy or decouple evidence acquisition from policy optimization via external modules, limiting their adaptability. In this paper, we propose **Route2Learn**, an end-to-end native agentic framework that learns query-adaptive routing between generation-based and retrieval-based strategies within a single video MLLM. Building on the insight that learning to route requires comparing both routes, we introduce a two-stage optimization approach: For Supervised Fine-Tuning, we actively execute both strategies for each query and progressively collect trajectories with rejection sampling, providing an implicit routing signal by exposing the model to which strategies succeed on different queries. We then introduce Counterfactual Routing Policy Optimization (CRPO), which explicitly optimizes routing preferences using the outcome difference between factual and counterfactual rollouts, while decoupling advantage estimation for tool selection and answer generation to enable precise credit assignment. Extensive experiments demonstrate that Route2Learn substantially outperforms existing strong baselines across four long-form video understanding benchmarks, improving over Qwen3-VL-8B by +10.4% on average. Further analysis shows that the model learns to select the more effective evidence-acquisition strategy according to the query. Our data, code, and model weights will be released publicly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.