acceptodds
Under review as a conference paper at ICLR 2027

Dissecting Reinforcement Learning: Mechanisms Behind Compositional Reasoning in LLMs

Abstract

Reinforcement learning with verifiable rewards (RLVR) has become the standard way to train large reasoning models, yet it is unclear which of its components account for its advantage over supervised fine-tuning (SFT). We study this question through the lens of compositional generalization: how well the model learns to compose multiple atomic skills to solve harder tasks. Our main contribution involves organizing post-training methods along two axes. The data axis classifies methods by the source of training data (teacher, frozen checkpoint, or on-policy models). The loss axis, in the binary reward setting, considers how different loss functions place weight on correct and incorrect samples during gradient updates. Through this unifying perspective, we are able to isolate the effect of on-policy data, negative gradients, the group-mean baseline, and standard-deviation normalization individually. On a string-function task and on competition math, we find that neither on-policy data nor negative gradients suffices alone. However, when combined, they improve average accuracy and allow the model to generalize beyond the compositional depths or difficulty of math on which it was trained. We provide a theoretical intuition for why this is the case: methods that give weight to negative examples allow the gradient to concentrate at critical points which distinguish successes from failures, while on-policy data ensures that these updates concentrate on wrong decisions the model would actually make. Consistent with this account, we find that training dynamics and results correspond to several predicted failure cases. Positive-only training with off-policy data such as teacher SFT is stable but lacks generalization, while on-policy positive-only training can lead to entropy collapse and shortened responses which eventually degrade performance. Lastly, applying negative rewards to off-policy data destabilizes training with large gradient norms and produces degenerate outputs. Combining on-policy sampling with negative gradients avoids these observed failure modes and supports sustained generalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.