GroundedVCoT: Grounding Interleaved Text-Image Reasoning with Reinforcement Learning
Abstract
Visual intelligence requires reasoning about how a scene changes and using the resulting state to guide later decisions. Unified multimodal models can make this explicit by interleaving textual reasoning with generated images. Yet generated images also enter the context for later steps. If any generated image misrepresents an action or contradicts the text, the error can compound across the trajectory. We introduce GroundedVCoT, a reinforcement-learning (RL) framework for grounding interleaved text–image chain-of-thought. At each step, task-specific verifiers assess textual claims and generated visual states against the state implied by the model's own decisions. Step-wise grounding rewards and task success yield a shared advantage that updates both text and image policies. Verification follows the model's trajectory, allowing alternative valid solutions while preserving its self-generated visual history. We evaluate three distinct demands on visual reasoning: planning moves in Rush Hour, constructing a semantic figure in Tangram, and tracking rotations in Rolling Dice. Relative to supervised fine-tuning, GroundedVCoT raises Rush Hour solve rate from 48.0% to 71.3% (+23.3%), Tangram accuracy from 25.6% to 50.5% (+24.9%), and Rolling Dice accuracy from 25.6% to 60.0% (+34.4%), also consistently exceeding terminal reward RL, supporting the value of grounding intermediate steps rather than rewarding task completion alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.