acceptodds
Under review as a conference paper at ICLR 2027

VisualRefineBench: Can Unified Multimodal Models Reliably Refine Their Own Images?

Abstract

Unified multimodal models (UMMs) can inspect and revise their own images at test time, but their gains in final image quality remain unclear. We introduce **VisualRefineBench**, combining end-to-end evaluation with three diagnostic tasks across 800 examples. Refine-E2E measures initial-to-delivered improvement and records the best image score during refinement. Under fixed inputs, Reflection tests error diagnosis, Decision tests action choice, and Editing measures net gains from instructed edits. With a four-image-call budget, **all six evaluated UMMs deliver lower mean quality than their corresponding Oracle best-of-N baselines**. Diagnostics reveal missed errors, poor action choices, and damage to correct content that often offsets repair gains. Replacing the Controller with a model stronger in Reflection improves mean final quality by **about 17 points** across four generator–editor combinations. These gains primarily reflect higher-quality candidates produced during refinement, while reliable final-image selection remains a challenge. These findings identify opportunities to improve UMM self-refinement and show how diagnostic evaluation can guide component choices that improve end-to-end outcomes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.