acceptodds
Under review as a conference paper at ICLR 2027

ReSynth: Verified Synthetic Data for Multi-Image Understanding

Abstract

People often understand the world by comparing different views of it. Vision-language models (VLMs), however, receive these views as separate images and still struggle to relate them. One problem is data. Training on these tasks requires groups of images with a known relationship: the same object from different viewpoints, or the same scene before and after a change. Such examples are hard to collect at scale from existing photo collections. We introduce ReSynth, which creates them entirely from generated images. Localized editing produces before-and-after pairs, viewpoint generation produces multiple views of the same instance, and verified scene contents support comparisons of counts and attributes. Generated images do not always match their prompts, so ReSynth does not assume that a requested relation was successfully rendered. Instead, it verifies facts such as object counts and whether an edit really happened, then builds questions only when the answer is supported by those verified facts. Using GRPO on 13K ReSynth questions, with no real images during training, Qwen2.5-VL-7B improves by 3.8 to 5.1 points on BLINK, MUIRBench, and VLM2-Bench, while single-image performance remains essentially unchanged. In a controlled comparison, using the object counts actually visible in generated images rather than the counts requested in their prompts improves counting accuracy by 9.75points. ReSynth suggests that generated images can provide useful multi-image supervision. The key is to verify what the generated images actually contain before using them for training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.