Reinforcing Diffusion Language Models with Clean State Objectives along Denoising Trajectories
Abstract
Policy loss estimation remains a fundamental and long-standing challenge in reinforcement learning (RL) for diffusion language models (DLMs). We introduce **Clean State Policy Optimization (CSPO)**, a novel training paradigm that selectively samples denoising timesteps and directly targets the final clean state for policy optimization. To reduce computational cost and stabilize training dynamics, CSPO optimizes the model toward the clipped clean state from intermediate noisy states, combined with weighted timestep sampling over denoising timesteps. Extensive experiments demonstrate that CSPO achieves consistent and substantial improvements in both performance and generalizability across two representative DLM architectures, LLaDA and Dream, on multiple reasoning benchmarks. Our work provides a practical approach to efficient and stable reinforcement learning in diffusion language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.