acceptodds
Under review as a conference paper at ICLR 2027

DPR: From High-Entropy Tokens to Learning Claims — Matched Evaluation for Token-Selective LLM Reinforcement Learning

Abstract

Decision-point reinforcement (DPR) is a verifier-free algorithmic framework for routing sequence-level reinforcement signals to informative response tokens. For each prompt, the policy generates on-policy rollouts and constructs a consensus pseudo-label from semantically equivalent parsed answers, yielding group-level reward signals. A frozen reference model ranks response positions by next-token entropy. DPR applies a clipped policy-gradient objective to the highest-entropy positions and anchors complementary tokens to the reference policy through a low-variance KL penalty. This dual-mask objective integrates entropy-guided credit assignment, reference-policy regularization, and total-token normalization within a unified training procedure. A count-matched random-routing counterpart and shared optimization contract enable precise evaluation of token-location quality. Controlled experiments show that DPR outperforms matched random routing in the anchor and data-shift evaluations, with favorable effects across learning checkpoints. Overall, DPR provides an effective and auditable approach to selective credit assignment for language-model reinforcement learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.