acceptodds
Under review as a conference paper at ICLR 2027

When Reflection Fails: Improving Self-Reflective Unified Multimodal Models with Correction Amplification and Fallback

Abstract

Self-reflective unified multimodal models (UMMs) offer a promising paradigm for complex image generation by enabling a single model to inspect its own outputs and iteratively refine them. However, they often struggle with generation tasks involving multiple objects and spatial relationships. To mitigate this problem, existing approaches typically rely on utilizing agents to perform reflection and generation separately, rather than exploiting the unified model's intrinsic capability. We revisit the source of these failures within the unified model and identify two distinct bottlenecks: reflection failure, where the model fails to diagnose errors in its reflection, and correction failure, where the model cannot effectively translate correction feedback into visual outputs. Based on this decomposition, we propose RECAF, a training-free framework that enhances both stages while retaining the unified model as the central component for reflection and generation. Experiments on GenEval, GenEval++, and GenEval2 demonstrate that RECAF consistently improves self-reflective image generation, outperforming baselines by 5.0, 5.7, and 3.5 points, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.