acceptodds
Under review as a conference paper at ICLR 2027

FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing

Abstract

Vision-Language Models (VLMs) have shown strong promise on Optical Character Recognition (OCR), yet the sheer number of visual tokens needed to encode dense documents makes inference prohibitively expensive. Existing methods like SnapKV evict visual tokens, which often fails on OCR tasks given the high information density of document images. We observe that, although document images appear globally dense and seemingly unprunable, a VLM's top-attended visual-token sets are line-local and drift gradually across decoding steps, much as a human reader fixates on successive words rather than on a whole page at once. Turning it into a method, we propose **FastOCR**, a training-free framework with two complementary modules: *Focal-Guided Pruning* identifies a small set of focal layers and selects the most task-relevant visual tokens at each step, while *Cross-Step Fixation Reuse* warm-starts each step from the previous one. FastOCR changes only which cached entries are attended and evicts none, so every token stays recoverable at every later step. Under matched nominal reference attention budgets, FastOCR outperforms every baseline we evaluate on OmniDocBench and olmOCR-Bench, and generalizes across various VLMs. On Qwen2.5-VL it retains **98.3%** of the unpruned model's accuracy on OmniDocBench and **97.3%** on olmOCR-Bench, with a **5%** visual-token budget at pruned layers. Latency measurements reach up to **3.4** attention speedup on olmOCR.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.