FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
Abstract
Vision-Language Models (VLMs) have shown strong promise on Optical Character Recognition (OCR), yet the sheer number of visual tokens needed to encode dense documents makes inference prohibitively expensive. Existing methods like SnapKV evict visual tokens, which often fails on OCR tasks given the high information density of document images. We observe that, although document images appear globally dense and seemingly unprunable, a VLM's top-attended visual-token sets are line-local and drift gradually across decoding steps, much as a human reader fixates on successive words rather than on a whole page at once. Turning it into a method, we propose **FastOCR**, a training-free framework with two complementary modules: *Focal-Guided Pruning* identifies a small set of focal layers and selects the most task-relevant visual tokens at each step, while *Cross-Step Fixation Reuse* warm-starts each step from the previous one. FastOCR changes only which cached entries are attended and evicts none, so every token stays recoverable at every later step. Under matched nominal reference attention budgets, FastOCR outperforms every baseline we evaluate on OmniDocBench and olmOCR-Bench, and generalizes across various VLMs. On Qwen2.5-VL it retains **98.3%** of the unpruned model's accuracy on OmniDocBench and **97.3%** on olmOCR-Bench, with a **5%** visual-token budget at pruned layers. Latency measurements reach up to **3.4** attention speedup on olmOCR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.