acceptodds
Under review as a conference paper at ICLR 2027

PeerCredit: Cross-Prompt Token Credit Assignment for Reinforcement Learning with Verifiable Rewards

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves reasoning models using automatically checkable outcome rewards, but its response-level advantages do not specify how negative credit should be distributed across tokens in failed responses. Existing methods derive finer-grained signals mainly from rollouts of the same prompt, leaving cross-prompt reward information unused. We propose PeerCredit, which introduces cross-prompt reward information as an additional source of evidence for negative credit assignment. PeerCredit constructs reward-weighted peer gradients, measures token-level gradient alignment, removes variation explained by token probability, relative token position, and response length, and redistributes a fixed negative-credit budget so that tokens better aligned with reward-preferred peer gradients receive less penalty while the total budget is preserved. Across eight benchmarks spanning mathematical reasoning and tool use, PeerCredit achieves the best mathematical average at all three Qwen3 scales, with gains of up to 15.0 points over GRPO and 2.7 points over ResRL, while also leading on both tool-use benchmarks. Further analyses support cross-prompt reward evidence as a useful principle for improving fine-grained credit assignment in RLVR beyond within-prompt supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.