Stochastic yet Structured: Reasoning Branch Selection Dynamics under RLVR
Abstract
Learning dynamics of LLM post-training provides insight into how model behaviors emerge and evolve throughout optimization. In this paper, we study the learning dynamics of reinforcement learning with verifiable rewards (RLVR) by tracking the evolution of probability distributions over behavioral branches along training trajectories and controlled continuations from fixed training states. By projecting RLVR training trajectories into an interpretable behavioral state space, we characterize RLVR’s reallocation of probability mass among behavioral branches across models and tasks. Controlled continuations from the same complete training state show that future branch allocations are not fully reproducible, even though the dominant branch often remains stable. Across continued training, these differences retain only short-range directional memory, while the current branch state provides a near-first-order predictive summary of short-horizon dynamics. The remaining variation is nevertheless structured: residual dependence shapes long-horizon divergence, while the current branch state predicts future instability across independent lineages and over substantially longer training horizons. Together, these results characterize RLVR branch selection as a stochastic but state-dependent dynamical process.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.