Belief-Shift Reward: Redistributing RLVR Credit with Policy-Internal Target Support
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) provides a scalable framework for improving reasoning post-training, but binary outcome rewards cannot distinguish among trajectories that an external verifier judges correct. To address this limitation, we propose Belief-Shift Reward (BSR), a lightweight reward transformation that preserves the external verifier as the sole criterion of correctness. Within the verifier-correct set, BSR differentiates trajectories by the change each one induces in the policy's support for the verified target relative to the prompt-only prior. Because this prior is shared within a prompt, it cancels under correct-set normalization. BSR therefore requires at most one detached posterior target-scoring pass per verifier-correct trajectory, without an auxiliary reward model or explicit step-level scoring. Across mathematical reasoning, agentic search, and code-output prediction, BSR yields aggregate improvements over matched RLVR baselines, spanning multiple model families and four group-relative optimization objectives. These results support the practical utility of policy-internal target support for finer-grained within-correct credit assignment when the verifier target can be scored by the policy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.