PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding
Abstract
Document and GUI understanding are among the highest-value applications of Vision-Language Models (VLMs), yet they impose exceptionally heavy computational burden: fine-grained text and small UI elements demand high-resolution inputs that produce tens of thousands of visual tokens. We observe that this cost is largely wasteful: across document and GUI benchmarks, 29–78% of patches within an image are pixel-level duplicates of others. We propose PixelPrune, which exploits this pixel-level redundancy through predictive-coding-based compression, pruning redundant patches before the Vision Transformer (ViT) encoder. Because it operates in pixel space prior to any neural computation, PixelPrune accelerates the entire inference pipeline—ViT encoder, patch merger, and LLM decoder—and applies equally to training. PixelPrune is parameter-free and supports pixel-lossless () as well as controlled lossy () compression. Without any training, it stays on average within 0.9 points of full-token accuracy on documents and similarly close on GUI question answering. On GUI grounding, where predicting precise coordinates makes the model more sensitive to the mismatch between full-token training and pruned inference, distillation from the model's own full-token outputs restores full-token accuracy. Trained from scratch on GUI grounding with PixelPrune, a model matches full-token training and can be served with either pruned or full-token input. PixelPrune speeds up inference by 3.0–4.2 and training by 1.9.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.