DUETTO: Dual-Channel Efficient Token Selection for Large Vision-Language Models
Abstract
Large Vision-Language Models introduce significant computational challenges, as encoding images and videos into thousands of vision tokens imposes substantial computational costs. Prior work addresses this challenge by proposing methods to identify and select the most informative vision tokens while preserving model performance. Determining which vision tokens are important requires considering two signals: the information a token contributes relative to the rest of the vision token sequence and its relevance to the instruction. Yet, existing methods fail to jointly optimize both for the downstream task: heuristic approaches rely on fixed ranking rules that cannot adapt to the downstream effects of token selection, while learnable approaches either base selection solely on visual content or defer instruction relevance to a later stage. To address this, we propose DUETTO (Dual-channel Efficient Token selection), a vision token selector trained end-to-end that fuses a vision channel and an instruction-relevance channel into a single selection score. DUETTO achieves state-of-the-art performance, with the advantage over prior methods increasing as the token budget becomes more restrictive, maintaining 99.1% of full-model accuracy on 33.3% of the vision tokens and, at high resolution, 96.5% using 5.6% of the vision sequence, accelerating prefill by 3.2x while the selection costs only 6% of prefill time. Finally, experimental results show that DUETTO generalizes to diverse backbones with both image and video inputs, achieving state-of-the-art results.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.