acceptodds
Under review as a conference paper at ICLR 2027

Rendering as Supervision: Exact-Geometry Synthetic Documents for Grounded Document Understanding

Abstract

Production document AI can read but cannot _point_: models answer questions yet emit no dependable evidence location, blocking auditable extraction and human review. On our probe, open vision–language bases at every scale (0.8B–9B) produce no parseable box at all, and the strongest frontier model is coordinate-frame-unreliable (0.157 [email protected] as instructed; 0.892 only under an oracle frame) — while the same models read text from _given_ regions at up to 0.86 hit rate. The missing box supervision is unaffordable to annotate and unreliable to OCR precisely where degradation strikes. Our position: _when you synthesize the degradation, you own its geometry_. Our engine persists the exact, invertible correspondence field of every degraded view (882 B/page, bit-reproducible), turning 126K rendered pages into 233K zero-error grounding labels. A dosed two-stage recipe confers coordinate emission on a 4B model (0.821 mIoU, generalizing across document classes, model families, and languages), improves real-document extraction by +14.8 exact-EM, and regresses no tracked public benchmark. Controlled studies then map _how_ label quality matters: i.i.d. label noise at OCR grade costs 3–11 mIoU (monotone in both domains; 3 seeds synthetic), whereas an equal-dose systematic bias is learned and fully correctable (0.526 raw, 0.841 after un-shifting — the exact-label anchor): exactness pays against the error structure OCR produces. A token-level equivariance loss on the same field fails for a measured reason: the warp is sub-resolution for the vision tokenizer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.