Learning When to Look: Adaptive Visual Computation for Video MLLMs
Abstract
Video multimodal large language models (MLLMs) process long visual sequences, yet fixed token budgets ignore variation in the evidence needed by different video–question pairs. We introduce Route2K, which organizes visual tokens into nested prefixes and uses a low-rank predictor initialized from one real seed state to choose a route without repeated full-decoder decisions. The route sets early decoder width before question-conditioned late pruning. On a frozen 1,000-question NExT-QA validation subset, Route2K reaches 71.7% accuracy at 1,599 executed visual token-layers, versus 71.4% at 2,324 for a measured FastVID operating point; narrow execution reaches 69.0%. On MVBench and TempCompass, it executes fewer visual token-layers than selected FastVID points with reported accuracy within one percentage point. Validation ablations support route-conditioned entry widths. These results show the potential of sample-conditioned routing to improve the accuracy–computation balance of video question answering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.