acceptodds
Under review as a conference paper at ICLR 2027

Imago: Studying the Online RL Recipe for Visual Generation

Abstract

We introduce Imago, an online reinforcement learning (RL) post-training recipe that matures pretrained text-to-image (T2I) generators into reliable compositional models. While modern T2I generators possess rich visual representations, under compositional constraints they frequently omit entities, bind attributes incorrectly, and violate spatial relations. RL offers a principled framework to resolve this gap, yet prior research focuses primarily on optimization algorithms, leaving the reward signal and the prompt corpus comparatively underexplored. Evaluating candidate rewards on optimization stability and visual fidelity, we find that a zero-shot VQAScore evaluated by a multimodal large language model (MLLM) provides reliable guidance, whereas preference-trained rewards are less reliable, with some triggering reward hacking. Treating the prompt corpus as the exploration environment, we further show that short and long prompts develop complementary spatial and attribute capabilities. While naive corpus mixing degrades performance, a sequential short-to-long curriculum successfully captures the strengths of both distributions. We apply the resulting recipe to FLUX.2-Klein-Base and Z-Image without architectural modifications, producing Imago-4B and Imago-6B. Both models achieve leading performance in their respective size classes across compositional capability and visual appeal, with Imago-6B surpassing substantially larger open-source and proprietary models on medium- and long-prompt composition benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.