Beyond Verbal Reasoning: Visual Structure Reasoning Language for Image Generation in Unified Multimodal Models
Abstract
Recent unified multimodal models (UMMs) have begun to introduce explicit natural-language reasoning before image generation. We revisit this default choice and propose Visual Structure Reasoning Language (VSRL), a structured reasoning language for image generation in UMMs. Given a prompt, the UMM autonomously generates VSRL comprising the global scene, objects, and their relations, and conditions image generation on it. We develop a VSRL data construction pipeline that converts paired prompts and target images into verified prompt–VSRL–image triplets. The resulting data are used to train the UMM to generate and use VSRL through supervised fine-tuning. We further introduce StruFlow-GRPO, a reinforcement learning framework that jointly optimizes reasoning and generation with programmatic structure rewards and image rewards. On BAGEL, VSRL outperforms both the base model and BAGEL’s native natural-language reasoning (Self CoT) across five generation benchmarks. Compared with Self CoT, VSRL improves GenEval Short from 79.0 to 88.9 and OneIG-Bench from 32.4 to 43.7. On Cheers, we train two separate models on VSRL and natural-language CoT data for a controlled comparison. VSRL outperforms both natural-language CoT and the base model across all evaluated generation benchmarks, including GenEval and TIIF-Bench. These results highlight reasoning representation as an important design dimension for UMM image generation and the potential of generation-oriented structured reasoning across UMMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.