acceptodds
Under review as a conference paper at ICLR 2027

DuetGRPO: Co-Adaptive Joint Reinforcement Learning for Reasoning-Driven Image Generation

Abstract

Joint training allows the language and visual policies in reasoning-driven image generation to co-adapt prompt formulation and image synthesis. However, existing joint training faces a central challenge: relying solely on image rewards induces reward mismatch, as rendering quality is an unreliable proxy for reasoning and prompt correctness. In this paper, we propose DuetGRPO a joint reinforcement learning framework that combines modality-specific optimization with shared alignment feedback to support online co-adaptation of language and visual policies. To supervise language outputs directly, we use a text-only evaluator to assess reasoning correctness and check generation prompts against a candidate-independent checklist of instruction requirements. We route this text reward exclusively to the language policy and share image–text alignment feedback across both policies, aligning their joint updates with distinct objectives: the language policy learns to reason correctly and formulate prompts that satisfy the user instruction and can be faithfully rendered, while the visual policy learns to follow those prompts. Experiments improve WISE Overall from 0.72 to 0.77 and T2I-CoReBench Overall from 59.1 to 60.7. Improvements are also observed when this recipe transfers from modular MLLM + diffusion transformer (DiT) pipelines to a mixture-of-transformers (MoT) architecture.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.