acceptodds
Under review as a conference paper at ICLR 2027

SQ-Pruner: Spectral Saliency and Query-Guided Visual Token Pruning in Large Vision-Language Models

Abstract

Redundant visual tokens impose substantial computational overhead on Large Vision-Language Models (LVLMs). Prior studies typically estimate token importance from visual representations after patchification or spatial aggregation, potentially weakening fine-grained structural information. Moreover, existing query-guided pruning methods often rely on a single relevance signal, which may not fully capture the correspondence between visual tokens and the input query. To address these issues, we propose Spectral and Query-Guided Pruner (SQ-Pruner), a training-free visual token pruning framework that adopts a two-stage pipeline: 1) Spectral Saliency Anchoring (SSA) estimates pixel-level spectral saliency before visual encoding to preserve structurally informative tokens as anchors. 2) Query-Guided Relevance Selection (QRS) adaptively fuses query-related signals to select visual tokens from the remaining candidates. Extensive experiments demonstrate the effectiveness of our framework across diverse LVLMs and tasks in both image and video modalities. On Qwen2.5-VL, SQ-Pruner retains 81.6% of the original performance while pruning 95.1% of visual tokens, outperforming the strongest prior method by 4.3%. SQ-Pruner also achieves a 1.42 prefilling speedup over the original model, demonstrating a favorable accuracy-efficiency trade-off for LVLM inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.