ResAdapt: Adaptive Resolution for Efficient Video Reasoning
Abstract
Multimodal Large Language Models (MLLMs) typically encode video frames at uniform spatial resolution, forcing fine visual detail and temporal sampling density into a rigid trade-off under constrained token budgets. Yet, visual relevance in video is temporally sparse and query-dependent: critical evidence concentrates in brief moments, while repetitive background motion carries little useful signal. We introduce , an input-side framework that treats per-frame resolution as an adaptive variable determined prior to visual feature extraction. A lightweight Allocator pairs coarse preview features with the user query to predict continuous per-frame scales, coarsening redundant intervals while allocating denser token capacity to answer-critical moments. To optimize allocation without frame-level supervision under non-differentiable processor token collation, we formulate the task as a continuous-action contextual bandit. We train the Allocator via , combining group-relative advantage with dynamic-pivot asymmetric cost shaping, temporal-similarity regularization, and concentration penalties. Across diverse video benchmarks, ResAdapt improves the accuracy–budget frontier: on Qwen2.5-VL at 128 frames, it matches dense full-resolution accuracy on the six-benchmark average ( vs. ) at only token retention, while exceeding dense performance on VideoMME and VideoMMMU. Furthermore, spatial savings are fungible with temporal coverage—reinvesting tokens into denser sampling of the same video raises VideoMME by points under a matched token budget. Permutation controls show that deliberate resolution placement across frames contributes to these reasoning gains. Finally, the frozen Allocator transfers zero-shot across diverse backbones and cuts 128-frame latency by at retention on Qwen2.5-VL. Code is available at https://anonymous.4open.science/r/ResAdapt.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.