InterGRPO: Global-Local Credit Assignment for Multi-Unit Interleaved Text–Image Generation
Abstract
Multi-unit interleaved text-image generation aims to produce coherent sequences of alternating textual and visual content, such as visual storytelling generation, requiring semantic and visual consistency across units as well as coherence of the complete sequence. Yet models trained primarily through supervised fine-tuning have limited ability to improve these qualities from generation feedback. Although recent reinforcement learning methods offer a potential remedy, existing methods primarily optimize a single image or an atomic text-image unit. Extending them to multi-unit outputs poses two challenges: a trajectory-level reward cannot identify which visual units should be reinforced, and uncertain quality judgments can provide unreliable optimization signals.To address this challenge, we introduce InterGRPO, an online group-based policy optimization framework tailored to multi-unit interleaved text-image generation. InterGRPO employs global-local credit assignment to optimize individual visual units while preserving sequence-level coherence. However, scoring hallucinations is widespread in general vision-language models for image quality evaluation, leading to inaccurate quality scores. To address this issue, we further develop an Uncertainty-Calibrated Reward (UC-Reward) model, which distills calibrated score distributions from multiple task-specific teachers. Experiments on SIL-Bench and ISG-Bench show that InterGRPO outperforms the strongest baseline by 2.5% and 4.5%, respectively, while improving cross-unit consistency, logical coherence, and overall generation quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.