acceptodds
Under review as a conference paper at ICLR 2027

Reinforcing Diffusion Language Models with Clean State Objectives along Denoising Trajectories

Abstract

Policy loss estimation remains a fundamental and long-standing challenge in reinforcement learning (RL) for diffusion language models (DLMs). We introduce **Clean State Policy Optimization (CSPO)**, a novel training paradigm that selectively samples denoising timesteps and directly targets the final clean state for policy optimization. To reduce computational cost and stabilize training dynamics, CSPO optimizes the model toward the clipped clean state from intermediate noisy states, combined with weighted timestep sampling over denoising timesteps. Extensive experiments demonstrate that CSPO achieves consistent and substantial improvements in both performance and generalizability across two representative DLM architectures, LLaDA and Dream, on multiple reasoning benchmarks. Our work provides a practical approach to efficient and stable reinforcement learning in diffusion language models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.