acceptodds
Under review as a conference paper at ICLR 2027

Agent-Level OCR: A Benchmark and An OCR-Enhanced Agent

Abstract

Optical character recognition (OCR) has advanced from transcribing individual characters to recognition and layout-aware parsing, supported by increasingly capable models and benchmarks. Existing benchmarks, however, primarily evaluate a model response to a prescribed image or document. A complete agent faces a broader problem: it must recognize when OCR is needed, locate the relevant visual evidence, verify uncertain observations, and carry the recovered information to complete a task. We define this ability to acquire, verify, and use visual-text information during task execution as agent-level OCR capability. To measure it, we introduce OCR Agent Bench, which extends the evaluation unit from an isolated model response to a complete agent execution spanning the reasoning model, harness, skills, tools, and OCR backend. The benchmark contains 260 executable tasks that evaluate delivered artifacts against task-specific criteria and source evidence. This evaluation target also motivates CLOCR Agent. Pre-extracted OCR alone does not tell an agent where to inspect, when to obtain another observation, or how to associate recovered information with its source. CLOCR Agent instead turns OCR from fixed preprocessing into an on-demand, locally verifiable, and source-associated capability through OCR-oriented guidance, source-addressed evidence operations, structured content access, and selective local OCR. Across 17 evaluated model–harness configurations, CLOCR Agent achieves the highest overall score (68.3) and ranks first for three of the four reasoning models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.