InterleaveThinker: Reinforcing Agentic Interleaved Generation
Abstract
Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their architectures, they cannot achieve interleaved generation (text-image sequence), which has crucial applications in visual narratives, guidance, and embodied manipulation. Even the latest open-source Unified Multimodal Models (UMMs) exhibit limited performance in this regard. In this paper, we introduce InterleaveThinker, the first multi-agent pipeline designed to endow frozen image generators with interleaved generation capabilities. Specifically, we employ a planner agent to organize the image-text input sequence, instructing the image generator on the required execution at each step. Subsequently, we introduce a critic agent to evaluate the generator's outputs, identify samples that deviate from the planned instructions, and refine the instructions for regeneration. To implement this pipeline, we construct the Interleave-Planner-SFT-80k and Interleave-Critic-SFT-112k to perform a format cold-start. Then we develop Interleave-Critic-RL-13k to reinforce the instruction correction capability within a generation trajectory using GRPO. Since an interleaved generation trajectory may involve over 25 generator calls, full-trajectory optimization is computationally expensive and suffers from difficult credit assignment. We therefore propose **T**rajectory-**D**ecoupled **GRPO** (**TD-GRPO**), which performs single-step optimization with accuracy and step-wise rewards to provide effective learning signals throughout the generation process. InterleaveThinker shows consistent gains across two off-the-shelf image generators and achieves performance comparable to Nano Banana and GPT-5 on interleaved generation benchmarks. Beyond generation, it also substantially improves the base model on reasoning-based benchmarks; for example, on 4-step FLUX.2-klein, WISE increases from 0.47 to 0.73 and RISE from 13.3 to 28.9.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.