acceptodds
Under review as a conference paper at ICLR 2027

R3: Rubric-Guided Region Reward for Precise Text-to-Image Generation

Abstract

Modern text-to-image (T2I) models support complex, structured conditions involving multiple entities, attributes, text, and spatial layouts, yet existing image rewards remain insufficiently granular. Global rewards yield a single score for the entire image, and recent rubric-based rewards improve semantic granularity but judge each criterion against the whole image. We introduce **R3** (**R**ubric-guided **R**egion **R**eward), which reformulates image evaluation as an explicit target–evidence matching problem. R3 first structures the generation intent into entity-bound, coarse-to-fine rubrics and unambiguous grounding labels; then independently records what is present in the image through visual observations; finally, it matches visual evidence with the target specification while explicitly accounting for the spatial consistency of the observed targets. This design turns the reward estimation from direct holistic judgment into grounded verification: structured intent in, grounded credit out. We also introduce **GranuBench**, featuring high-quality contemporary generations, complex structured specifications, and explicit spatial targets to systematically evaluate reward granularity. On GranuBench, R3-based Gemini reaches a preference accuracy of 90.8%, outperforming direct-Gemini (80.7%) and existing reward models by a substantial margin. We further construct **R3Traj-100K**, containing 100K high-quality multi-stage R3 supervision trajectories, each recording structured target specifications, localized visual observations, and target–evidence assessments. Using R3Traj-100K, we distill the Gemini-instantiated R3 pipeline into an open 27B VLM with stage-specialized adapters, achieving state-of-the-art performance on multiple reward benchmarks. R3 can also improve image quality at test time by enabling precise region-targeted refinement. Together, these results demonstrate that finer reward granularity benefits not only evaluation but also inference- and training-time optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.