acceptodds
Under review as a conference paper at ICLR 2027

TRPQuant: Token Rearrangement and Channel Pruning for Quantizing Large Vision-Language Models

Abstract

High memory and computation costs prevent the practical edge deployment of vision-language models (VLMs). Post-training quantization (PTQ) offers a promising solution without expensive retraining. Given the visual and textual inputs in VLMs, existing methods typically adopt modality-level token partitioning and processing. However, they treat tokens uniformly within each modality, overlooking intra-modality activation distribution differences. This coarse-grained strategy leads to degraded accuracy and redundant computation under low-bit quantization. We propose TRPQuant, a new framework that addresses these limitations through Token Rearrangement and Channel Pruning, enabling fine-grained token grouping and computation reduction. Specifically, TRPQuant classifies tokens into Dense, Intermediate, and Sparse groups based on their activation patterns and applies tailored quantization strategies accordingly: dense tokens preserve full channels and go through the standard branch, sparse tokens are routed to the sparse branch with channel pruning to reduce computation, and intermediate tokens combine both mechanisms via dual-branch processing. Extensive experiments on mainstream VLMs demonstrate that TRPQuant achieves strong quantization accuracy. Notably, TRPQuant preserves 94.1% of FP16 performance under the challenging W3A3 setting (69.9 vs. 74.3) on Qwen2.5-VL-7B, pushing the frontier for ultra-low-bit quantization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.