acceptodds
Under review as a conference paper at ICLR 2027

ReflowDiff: Can Vision-Language Models Reliably Compare Documents Across Layouts?

Abstract

Vision-language models (VLMs) increasingly read, write, and revise documents. Yet it remains unclear whether they can report every textual revision when a document's layout also changes. Concretely, one changed digit or sign alters what a contract or financial report says, and catching it requires fine-grained character recognition, matching corresponding text across layouts, and exhaustive reporting. We introduce ReflowDiff, a benchmark of 4,488 comparison tasks built from 374 dense rich-text pages spanning six generated domains, real web pages, and financial reports. Each page is rendered in two layouts that preserve its text and structural context, and the same set of one to five character edits is applied to both. Two no-edit controls, Identical and Reflow only, test whether models report edits when the text is unchanged. We leverage exact match as the primary metric, which requires every reference edit, with exact before and after text at the correct occurrence, and no additional edits. Across 30 configurations of 26 VLMs, the best reaches 76.4% exact match when edits are combined with reflow. Reflow lowers exact match on the same edits in all 30 configurations, including a 3.8-percentage-point drop for the strongest, and elicits false reports even from configurations that make none on identical pairs. Even when models recover most individual edits, their reports remain incomplete and vary across layouts. These failures highlight a gap between reading document content and reliably verifying what has changed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.