Reflection Rewards Pay for Deferral: Capability-Aware Rewards for Reflective RL
Abstract
Self-reflective reinforcement learning trains models to revise their answers and is typically judged by final-answer accuracy. However, a reward for correcting mistakes also shapes the answer being corrected. We identify deferral, a failure mode in which a model postpones its real answer until after reflection, so that competitive final accuracy conceals a degraded first answer. In a controlled setting, a reflection-trained model's first answer falls below chance while its final answer remains accurate. We trace this to the reward: when a trajectory-level reward values a correction above a preserved correct answer, it introduces a bias toward deferral that grows with the model's capability. Natural remedies do not resolve this. Rewarding the reflection or answer changes can also pay for deferral, and since a deferred first answer is predictably wrong, rewarding calibrated confidence reinforces it too. We therefore propose Uncertainty-Aware Self-reflective Reinforcement Learning (UASRL), which uses a per-prompt capability proxy estimated once before RL, beyond the policy's control, and favors correct first answers where capability is high while still rewarding correction. Added to an existing reflection reward, UASRL substantially improves the first-answer accuracy of that reward and reduces deferral while maintaining competitive final accuracy. Across two backbones and three benchmarks, including a held-out one, it outperforms all tested reflection methods on first answers. Our results expose first-answer degradation as a hidden cost of reflection rewards, derive design criteria for such rewards, and show that self-reflection should be evaluated by both its first and final answers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.