acceptodds
Under review as a conference paper at ICLR 2027

STP: Efficient Reinforcement Learning for Diffusion Large Language Models

Abstract

Reinforcement Learning (RL) is crucial for unlocking the complex reasoning capabilities of Diffusion-based Large Language Models (dLLMs). However, applying RL to dLLMs faces unique challenges in efficiency and stability. To address these challenges, we propose Spatio-Temporal Pruning (STP), a framework that removes redundancy along both dimensions of the diffusion process. STP employs static priors, anchors reliable positions, and concentrates exploration, while leveraging bidirectional masked generation to bypass redundant late-stage refinement. Theoretically, we show that STP reduces rollout sampling cost and characterize the estimation bias–variance trade-off of spatial pruning and the output perturbation induced by temporal pruning. We further relate the bias and variance of paired ELBO-difference estimates to error and variance bounds for GRPO objective. Extensive experiments demonstrate that STP outperforms the evaluated baselines in accuracy and RL training efficiency. Our code is available at https://anonymous.4open.science/r/STP-757E.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.