acceptodds
Under review as a conference paper at ICLR 2027

LaDi-RL: Latent Diffusion Reasoning Prevents Diversity Collapse in Reinforcement Learning

Abstract

Reinforcement learning has become a central paradigm for improving LLM reasoning, but most existing methods optimize policies over discrete token sequences. This creates a mismatch between the optimization space and the structure of reasoning: many important decisions are semantic, global, and trajectory-level rather than local token choices. Continuous latent-space RL offers a promising alternative, but the latent policy must model a complex, multi-modal distribution over valid reasoning trajectories. We therefore propose Latent Diffusion Reasoning with Reinforcement Learning (LaDi-RL), where a diffusion model generates latent reasoning trajectories through iterative denoising, enabling structured exploration and expressive distribution modeling. This comes with a credit-assignment challenge: the policy acts in latent space, but rewards are observed only after the latent is decoded into text, so a naive rollout cannot tell whether an incorrect answer comes from a poor latent trajectory or from an imperfect textual realization. We address this with hierarchical latent-text rollouts, which decode several answers per latent and score the latent by their mean reward, separating the quality of the reasoning from the luck of its decoding. Empirically, LaDi-RL outperforms token-level RL by 10.5% on code generation and 6.5% on math reasoning in pass@1, and raises pass@k above its pre-RL starting point, which token-level RL fails to do.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.