Focus on What Matters: Text-Aware, Training-Free Token Pruning for Efficient Document Parsing
Abstract
Document parsing methods based on vision-language models (VLMs) typically encode high-resolution document images into dense sequences of visual tokens, resulting in high computational costs. Many of these tokens correspond to blank or background regions, resulting in redundant computation. Token pruning offers an effective way to alleviate this overhead by reducing the number of visual tokens. However, general-purpose token pruning methods suffer from two main limitations: (1) they typically rank visual tokens based on attention scores or feature similarity, potentially discarding tokens from critical text regions in the image; and (2) they struggle to adapt pruning ratios to varying text densities. Both limitations degrade parsing accuracy. To address these limitations, we propose -Doc, a Text-aware, Training-free Token pruning framework for efficient Document parsing. -Doc employs an off-the-shelf text detector to locate text regions and maps them onto the VLM's token grid to construct a retention mask. The framework incorporates two key designs: (1) Flexible Pruning Placement, supporting pruning before or after the visual encoder, with pre-encoder pruning further reducing the encoder's computational cost; and (2) Coverage-Adaptive Thresholding, which uses region dilation and bounding-box merging to regulate mask coverage, thereby adaptively adjusting the token retention ratio. Extensive experiments demonstrate that -Doc prunes approximately 30% of visual tokens on average while largely maintaining parsing accuracy, achieving a better accuracy–efficiency trade-off than existing methods. The code will be publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.