acceptodds
Under review as a conference paper at ICLR 2027

Efficient Reinforcement for Visual-Textual Thinking with Discrete Diffusion Models

Abstract

Despite promising multimodal understanding results, discrete diffusion models reason predominantly in text, using images as inputs but not as output for visual thinking that supports grounding and spatial reasoning. We introduce GRPO-based post-training that enables a unified discrete diffusion model to first generate an intermediate reasoning image and then use it to guide textual reasoning, while addressing two key challenges. First, to reduce visual rollout costs during RL, we leverage bidirectional attention for **localized visual editing**, denoising only a small subset of tokens while preserving all the other tokens. Second, we introduce **modality-factorized credit assignment** with segment-specific rewards, ensuring that image updates exclude future text while text updates use the completed image. Across five multimodal benchmarks, **LocFac-RL** improves average accuracy over standard GRPO by on LaViDa-O and on MMaDA-Parallel. Compared with global editing on the same backbones, it reduces training time by and while improving average accuracy by and , respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.