acceptodds
Under review as a conference paper at ICLR 2027

On Delayed Learning in Reinforcement Learning with Verifiable Rewards: Provable Laws of Feedback and Recovery

Abstract

Reinforcement learning with verifiable rewards (RLVR) turns automatically checked outcomes into improvements in reasoning policies. Yet informative updates can coexist with prolonged learning delays, even when the training run eventually reaches its target. To explain this phenomenon, we develop a finite-sample theory of feedback arrival and the training time needed to reach a fixed accuracy in RLVR. Our theory identifies a feedback mechanism that amplifies temporary regressions: when successful answers are rare, a harmful update makes the next informative sample harder to obtain, converting short sequences of regressions into heavy-tailed training costs. The severity of this effect is determined jointly by feedback scarcity and the temporal dependence of conditional progress. For a signed-update model, an exact dependence threshold characterizes when the expected training budget is finite. Stronger persistence can cross this threshold while preserving the stationary distribution of individual updates, producing infinite mean activation time at a fixed initialization and fixed accuracy target. Theoretically, we construct exponentially mixing processes with identical finite-window update distributions but different mean-cost regimes, and establish finite-initialization coupling and full-state bounds. We validate the predicted mechanisms empirically via controlled stochastic experiments and language model experiments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.