ReflowDiff: Can Vision-Language Models Reliably Compare Documents Across Layouts?
Abstract
Vision-language models (VLMs) increasingly read, write, and revise documents. Yet it remains unclear whether they can report every textual revision when a document's layout also changes. Concretely, one changed digit or sign alters what a contract or financial report says, and catching it requires fine-grained character recognition, matching corresponding text across layouts, and exhaustive reporting. We introduce ReflowDiff, a benchmark of 4,488 comparison tasks built from 374 dense rich-text pages spanning six generated domains, real web pages, and financial reports. Each page is rendered in two layouts that preserve its text and structural context, and the same set of one to five character edits is applied to both. Two no-edit controls, Identical and Reflow only, test whether models report edits when the text is unchanged. We leverage exact match as the primary metric, which requires every reference edit, with exact before and after text at the correct occurrence, and no additional edits. Across 30 configurations of 26 VLMs, the best reaches 76.4% exact match when edits are combined with reflow. Reflow lowers exact match on the same edits in all 30 configurations, including a 3.8-percentage-point drop for the strongest, and elicits false reports even from configurations that make none on identical pairs. Even when models recover most individual edits, their reports remain incomplete and vary across layouts. These failures highlight a gap between reading document content and reliably verifying what has changed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.