acceptodds
Under review as a conference paper at ICLR 2027

Beyond Reward Sparsity: How Sample Difficulty Shapes Reasoning in RLVR

Abstract

Reinforcement learning with verifiable rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large language models, yet how training sample difficulty shapes the learning process remains poorly understood. Existing analyses often view difficulty through the lens of reward sparsity, but reward availability alone does not explain why some samples improve reasoning while others lead to unstable updates. In this work, we present a learning-dynamics analysis of sample difficulty in RLVR. We distinguish static sample difficulty, measured by base-policy solvability, from dynamic learning-signal availability under the evolving policy, and derive a per-update decomposition of RLVR optimization into three components: the training signal, its transfer through shared parameters, and the resulting prediction change. Through controlled experiments on MATH, we show that different difficulty regimes induce distinct learning behaviors: frontier samples provide stable improvements, while hard samples can produce harmful updates despite retaining non-trivial signals. Using sparse autoencoders and feature interventions, we further reveal that difficulty-dependent training reshapes internal feature dynamics, with samples of similar difficulty producing heterogeneous feature changes depending on their realized rollout groups. These results provide a mechanistic understanding of how sample difficulty governs RLVR learning dynamics beyond reward sparsity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.