acceptodds
Under review as a conference paper at ICLR 2027

Layout-Aware Latent Evidence Reasoning for OCR-Free Key Information Extraction

Abstract

Key information extraction (KIE) requires associating textual values with their spatial context and requested fields. OCR-free models typically learn these associations implicitly from answer supervision, while OCR-augmented methods expose explicit evidence at the cost of external OCR dependence and variable-length context at inference. We propose **LAYER**, a **LAY**out-aware latent **E**vidence **R**easoning framework that uses OCR as training supervision to learn a compact internal representation of document evidence. Given a document image and a field query, LAYER predicts a fixed-length sequence of continuous visual latent tokens in parallel. The same language model then conditions on these latents and the original image-query context to autoregressively generate structured answers. Two complementary objectives guide learning: Layout-aware Feature Alignment aligns the predicted latents with visual features of OCR text rendered in its page layout, while Confidence-Gated Contrastive Learning associates answer-value representations with supporting OCR blocks through reliable field-region matches. A two-stage training scheme initializes the latent interface and then jointly optimizes structured generation with both objectives as auxiliary supervision. Inference requires only the document image and query. Experiments on public KIE benchmarks show pooled micro-F1 gains of 1.66%-3.45% over answer-only fine-tuning across three backbone sizes. A controlled timing comparison further shows a end-to-end speedup over OCR as Chain-of-Thought, with answer length fixed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.