Pack, Don't Drop: Breaking the Selection Bottleneck in Visual Token Pruning
Abstract
Visual token pruning reduces the inference cost of multimodal large language models (MLLMs), yet under strict token budgets, selection alone inevitably discards complementary visual evidence. Existing aggregation methods reuse information from pruned tokens within compressed sequences, providing a complementary direction beyond selection. Under a fixed token budget, effective information reuse therefore requires both appropriate routing of pruned content and suitable representations at the retained positions. To this end, we propose PackAdapt, which enhances a given set of retained positions without changing the final token budget. Its Information Folding module learns task-relevant routing from pruned tokens to retained tokens on top of semantic similarity, while its Representation Adaptation module further transforms the folded features through a lightweight residual mapping. With the base MLLM frozen, the selector, routing module, and representation adapter are jointly trained with the language-modeling objective, so supervision from the final prediction reaches all three components. Extensive experiments across multiple multimodal backbones, image-understanding benchmarks, and video-understanding benchmarks demonstrate that PackAdapt consistently improves compressed-model performance under different token budgets, with particularly strong gains under aggressive visual compression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.