acceptodds
Under review as a conference paper at ICLR 2027

Riding the Waves of Uncertainty: Distributional Value Learning for LLMs with Verifiable Rewards

Abstract

Reinforcement learning with verifiable rewards (RLVR) trains large language models from feedback on completed responses, yet learning a useful value baseline from sparse terminal outcomes remains difficult. A conventional PPO critic regresses a scalar expected return; a multi-output value objective offers an alternative way to learn from the same feedback. We propose DistRLVR, a distributional actor–critic framework that predicts prefix-conditioned returns and uses their expectation as the PPO baseline. Dual Sample Replacement (dSR) constructs multi-step critic targets from existing trajectories and target-critic predictions, supplying additional supervision without generating additional responses. The predicted quantiles also support optional tail-aware advantage reweighting. Experiments on mathematical reasoning and executable code generation show improvements over PPO and GRPO, including gains over VAPO in mean code reward; instruction-following results further examine transfer to graded programmatic feedback. Scalar-target and capacity controls support an optimization-based interpretation of the gains: multi-output losses can improve value learning without requiring accurate reconstruction of the complete return law. DistRLVR with dSR provides an effective RLVR training recipe, while tail reweighting yields a task-dependent trade-off between average quality and repeated-sampling success. Code is available in the anonymous repository (https://anonymous.4open.science/r/DistRLVR_Anonymous-CD2E/).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.