acceptodds
Under review as a conference paper at ICLR 2027

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Abstract

Unified multimodal models can both understand and generate images, making it possible, in principle, for them to iteratively repair their own generations: diagnose what an image gets wrong, revise it, observe the rendered result, and reflect again. Because whether a revision is effective can only be determined after rendering, reflection and image generation must be optimized jointly over the entire interaction trajectory. Supervised fine-tuning (SFT) on reflection trajectories teaches the interleaved reasoning-and-generation format, but does not distinguish which reflections actually lead to better images; meanwhile, naive reinforcement learning (RL) that optimizes only the renderer or a single output head leaves much of the potential gain untapped. We introduce **UMM-Reflection**, a trajectory-level RL framework that optimizes complete reflection trajectories within a single unified multimodal model. Sibling trajectories share the same initial image, allowing group-relative advantages to directly compare alternative reflection strategies, while a single trajectory-level advantage jointly updates both the reflection tokens and the flow-based image revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing methods or agentic pipelines that rely on an external critic, UMM-Reflection propagates credit across multiple rounds and across both roles of the same model, while requiring no external verifier at inference time. Built on BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, with strong zero-shot transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used during training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.