acceptodds
Under review as a conference paper at ICLR 2027

Beyond Verbal Reasoning: Visual Structure Reasoning Language for Image Generation in Unified Multimodal Models

Abstract

Recent unified multimodal models (UMMs) have begun to introduce explicit natural-language reasoning before image generation. We revisit this default choice and propose Visual Structure Reasoning Language (VSRL), a structured reasoning language for image generation in UMMs. Given a prompt, the UMM autonomously generates VSRL comprising the global scene, objects, and their relations, and conditions image generation on it. We develop a VSRL data construction pipeline that converts paired prompts and target images into verified prompt–VSRL–image triplets. The resulting data are used to train the UMM to generate and use VSRL through supervised fine-tuning. We further introduce StruFlow-GRPO, a reinforcement learning framework that jointly optimizes reasoning and generation with programmatic structure rewards and image rewards. On BAGEL, VSRL outperforms both the base model and BAGEL’s native natural-language reasoning (Self CoT) across five generation benchmarks. Compared with Self CoT, VSRL improves GenEval Short from 79.0 to 88.9 and OneIG-Bench from 32.4 to 43.7. On Cheers, we train two separate models on VSRL and natural-language CoT data for a controlled comparison. VSRL outperforms both natural-language CoT and the base model across all evaluated generation benchmarks, including GenEval and TIIF-Bench. These results highlight reasoning representation as an important design dimension for UMM image generation and the potential of generation-oriented structured reasoning across UMMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.