Tokenizing the Unseen: Compact Visual Prefix Cache for Overcoming VLM Input Bandwidth
Abstract
Fine-grained visual understanding faces an input-bandwidth tradeoff: low-resolution inputs discard small details such as text rendered at a small size, while high-resolution inputs occupy hundreds of decoder positions, a burden that compounds when the same image must answer many questions. Existing approaches each pay a different cost: agentic methods search, zoom, or crop for every query, so latency grows with the number of questions; raising the input resolution inflates the visual sequence; and token-reduction methods lose much of their accuracy under aggressive compression. We propose a compact visual prefix that decouples a high-resolution branch from the visual input budget of a frozen language decoder. This branch reads the image at full resolution and compresses it into the prefix, so the decoder needs only a fixed number of visual positions, while the pretrained backbone and its native path are left unchanged. Computed from the image alone, the prefix can be cached once and reused across questions. Under a budget of 24 decoder positions, the learned prefix reaches 0.386 TextOCR word F1 against 0.154 for the strongest of five tested token reducers. On DocVQA, a single prefix per page serves all 4.68 questions on average. These results show that a compact visual prefix can extract and compress useful information from a high-resolution image, and that the result is reusable across questions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.