Efficient Reinforcement for Visual-Textual Thinking with Discrete Diffusion Models
Abstract
Despite promising multimodal understanding results, discrete diffusion models reason predominantly in text, using images as inputs but not as output for visual thinking that supports grounding and spatial reasoning. We introduce GRPO-based post-training that enables a unified discrete diffusion model to first generate an intermediate reasoning image and then use it to guide textual reasoning, while addressing two key challenges. First, to reduce visual rollout costs during RL, we leverage bidirectional attention for **localized visual editing**, denoising only a small subset of tokens while preserving all the other tokens. Second, we introduce **modality-factorized credit assignment** with segment-specific rewards, ensuring that image updates exclude future text while text updates use the completed image. Across five multimodal benchmarks, **LocFac-RL** improves average accuracy over standard GRPO by on LaViDa-O and on MMaDA-Parallel. Compared with global editing on the same backbones, it reduces training time by and while improving average accuracy by and , respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.