acceptodds
Under review as a conference paper at ICLR 2027

DiffBench: How Well Do MLLMs Understand Visual Differences?

Abstract

Vision-to-code tasks aim to reconstruct visual inputs into editable and executable artifacts with high visual fidelity. During reconstruction, visual difference detection is a key tool for locating reconstruction errors and guiding subsequent refinement. It makes evaluation of visual difference detection increasingly important. Existing difference detection benchmarks largely focus on synthetic artifacts or object-centric natural images, leaving complex documents in real-world underexplored. To address this gap, we introduce DiffBench, a high-quality benchmark of visual differences detection, featuring non-synthetic discrepancies in diverse real-world document types such as slides, magazines and reports. Besides, we propose an automated evaluation framework for assessing visual difference predictions across four dimensions including layout, relation, content integrity and visual properties. Our evaluation reveals that GPT-6 Astra achieves the strongest overall performance and is better than GPT-5.6 Sol, while Kimi K3 leads among open-weight models including its previous version Kimi K2.5. Across models, content integrity discrepancies like missing text and images are generally easier to detect, whereas precise alignment, visual properties, and multi-region layout changes remain substantially more challenging. Performance further drops on highly similar cases with subtle differences, where models increasingly miss true discrepancies. Overall, DiffBench remains challenging for current MLLMs while effectively capturing progress across successive model generations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.