acceptodds
Under review as a conference paper at ICLR 2027

Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts

Abstract

Precise instruction following in image generation, such as satisfying object counts and spatial relations, is optimized using learned evaluators such as object detectors and vision-language models that are unreliable themselves. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, reveal a capability gap in current models and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBENCH, with 10,000 tasks over 46 constraint types, and VVRBENCH-Challenge, with 720 more complex tasks; the strongest model we evaluate—GPT-Image-2.5—solves 21.4% of VVRBENCH-Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBENCH from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.