acceptodds
Under review as a conference paper at ICLR 2027

Learning When to Look: Adaptive Visual Computation for Video MLLMs

Abstract

Video multimodal large language models (MLLMs) process long visual sequences, yet fixed token budgets ignore variation in the evidence needed by different video–question pairs. We introduce Route2K, which organizes visual tokens into nested prefixes and uses a low-rank predictor initialized from one real seed state to choose a route without repeated full-decoder decisions. The route sets early decoder width before question-conditioned late pruning. On a frozen 1,000-question NExT-QA validation subset, Route2K reaches 71.7% accuracy at 1,599 executed visual token-layers, versus 71.4% at 2,324 for a measured FastVID operating point; narrow execution reaches 69.0%. On MVBench and TempCompass, it executes fewer visual token-layers than selected FastVID points with reported accuracy within one percentage point. Validation ablations support route-conditioned entry widths. These results show the potential of sample-conditioned routing to improve the accuracy–computation balance of video question answering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.