acceptodds
Under review as a conference paper at ICLR 2027

VDiff-Bench: Diagnosing and Improving Fine-Grained Visual Difference Understanding

Abstract

Multimodal large language models (MLLMs) achieve strong performance on general visual understanding tasks, yet reliably identifying differences between similar images remains challenging. We introduce VDiff-Bench, a diagnostic benchmark that connects controlled difference identification with open-ended difference discovery through two tracks: VDiff-Bench-MCQ, which consists of 1,756 multiple-choice questions across 10 semantic, textual, and low-level change categories; and VDiff-Bench-Gen, with 300 challenging image pairs for open-ended image difference generation. Evaluation of 12 MLLMs on the MCQ track reveals pronounced category-specific weaknesses, with small open-weight models particularly struggling on low-level changes. To address these failures without updating model weights, we propose Expert-guided Harness Evolution (EHE): an expert coding agent uses traces and evaluation feedback to revise the tools, routing, and response handling around a small base agent, successfully improving its MCQ performance by 26.9%, comparable with the GPT-5.4 model. To better evaluate open-ended image difference captioning, we further introduce VDiff-Judge, a tool-augmented agentic evaluator that judges generated differences using targeted visual evidence, forming verifiable traces for assessing correctness and completeness. Even the strongest reported proprietary model achieves only 46.8% micro-F1 on the Gen track. EHE improves Qwen3-VL-8B-Thinking from 65.7% to 83.4% MCQ accuracy. Notably, EHE's performance gain is transferrable to different data and different task: adapting EHE to the Gen track improves micro-F1 by 155.5%, approaching proprietary performance; adapting EHE to a different benchmark also yields over 90% improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.