Principled Reinforcement Learning with Verifiable Rewards: Understanding LLM Reasoning via Diffusive Flow
Abstract
Large language models (LLMs) increasingly rely on reinforcement learning with verifiable rewards (RLVR) for post-training on complex reasoning tasks. While most RLVR methods are policy-based, value-based approaches have shown promise but remain underexplored in stability, formulation, and theory. In this work, we revisit value-based RLVR through a smoothed consistency framework that stabilizes LLM fine-tuning. With reward augmentation, our framework generalizes the standard KL-regularized objective, connects value-based and policy-based methods, and yields a practical joint policy–value objective using only terminal verifiable rewards. To explain its effectiveness for LLM reasoning, we develop a reasoning-flow characterization over a reasoning-directed acyclic graph (DAG), showing how the learned RLVR objective induces latent stepwise credit assignment and reshapes probability mass over reasoning trajectories. This characterization reveals how the fine-tuned LLM favors high-reward reasoning paths while preserving diversity in structured reasoning. On the theoretical side, we establish a rigorous regret analysis showing that our algorithm achieves sample-complexity guarantees and is robust to the LLM policy drift. Experiments on mathematical reasoning benchmarks show consistent gains over pretrained LLMs and demonstrate that our algorithm is competitive with strong baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.