TokenGraft: Diagnosing and Improving Visual Token Compression via Independent Recovery
Abstract
Visual token compression reduces the visual sequence processed by vision–language models, but performance–budget curves leave one question unresolved: does a compressor use the available token slots effectively? At the same final token count, we compare two allocations of the available budget. One uses only compressor-produced tokens (native expansion), while the other combines fewer compressed tokens with selected pre-compression visual tokens (original-token recovery). Across multiple compressors, even simple recovery rules reveal a recoverable gap: this hybrid representation can outperform native expansion. We propose TokenGraft, a training-free, text-query-independent method that restores visually important original tokens from locations under-covered by the compressed representation. On Qwen2.5-VL, TokenGraft improves all 27 main grounding comparisons across three compressors, three retention rates, and RefCOCO, RefCOCO+, and RefCOCOg, with gains up to 33.57 percentage points. Fixed-total-budget sweeps further show how reallocating slots from compression to recovery improves accuracy. These gains extend to InternVL2.5, LLaVA-OneVision, and CV-Bench, with magnitudes that depend on the backbone, compressor, and task. Under our paired runtime protocol, recovery adds less than 1% full-request time across all three compressors. These results establish token allocation, alongside budget size, as a practical design choice for visual token compression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.