CODA: Contrastive Decision Advantages for Unified Text-to-Image Generation
Abstract
Unified models integrate language understanding and image synthesis within a single model, and reinforcement learning further improves their generation capability through outcome-level rewards on completed images. However, relying on a single final-image reward to jointly optimize semantic reasoning and visual generation leaves two difficulties: ambiguous semantic attribution and sparse visual supervision. The former mixes high-level semantic reasoning with its low-level visual realization, while the latter leaves a long chain of visual decisions without feedback before the image is completed. To address them, we propose **COntrastive Decision Advantag (CODA)**, a framework that derives contrastive decision advantages by comparing each decision with a matched alternative using shared randomness, which contains **Semantic Reasoning Guidance** and **Masked Trajectory Guidance**. Specifically, Semantic Reasoning Guidance isolates the effect of prompt optimization by comparing images generated from the optimized and original prompts. Masked Trajectory Guidance contrasts each probed commitment with equally sized alternatives and weights its outcome gain by the induced distribution shift, promoting earlier commitments that are both correct and structurally influential. A trajectory-level outcome advantage is further incorporated to provide a global signal. Grounded in the same constraint-anchored evaluation from the original instruction, the semantic advantage supervises the text tokens forming the optimized prompt, while the process advantage supervises each probed commitment. Extensive experiments on GenEval, T2I-CompBench, and WISE demonstrate that CODA achieves outstanding performance across tasks involving complex scene composition and knowledge-intensive image generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.