acceptodds
Under review as a conference paper at ICLR 2027

Stochastic yet Structured: Reasoning Branch Selection Dynamics under RLVR

Abstract

Learning dynamics of LLM post-training provides insight into how model behaviors emerge and evolve throughout optimization. In this paper, we study the learning dynamics of reinforcement learning with verifiable rewards (RLVR) by tracking the evolution of probability distributions over behavioral branches along training trajectories and controlled continuations from fixed training states. By projecting RLVR training trajectories into an interpretable behavioral state space, we characterize RLVR’s reallocation of probability mass among behavioral branches across models and tasks. Controlled continuations from the same complete training state show that future branch allocations are not fully reproducible, even though the dominant branch often remains stable. Across continued training, these differences retain only short-range directional memory, while the current branch state provides a near-first-order predictive summary of short-horizon dynamics. The remaining variation is nevertheless structured: residual dependence shapes long-horizon divergence, while the current branch state predicts future instability across independent lineages and over substantially longer training horizons. Together, these results characterize RLVR branch selection as a stochastic but state-dependent dynamical process.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.