acceptodds
Under review as a conference paper at ICLR 2027

CreditFilter: Calibrated Token Weights over the Within-Response Entropy Rank for RLVR

Abstract

Critic-free reinforcement learning with verifiable rewards (RLVR) typically assigns a uniform advantage to every token. Token-reweighting heuristics like the 80/20 mask favor high-entropy tokens, but fail to measure a token's actual impact on the outcome. We formalize this impact as the forking value: the variance of the success probability over the policy's next-token choices. Because the expected forking values along a sequence sum exactly to the reward variance, token credit acts as a budget each response spends at its own forks; consequently, weights should redistribute credit strictly within a response. Mathematically, the forking value factorizes into a fork consequence, which requires costly continuations to estimate, and a fork probability, which is intrinsically provided by the softmax. On Qwen3-8B-Base, this fork probability decays exponentially with a token's within-response entropy rank at class-specific rates, revealing that most tokens carry negligible forking value. CreditFilter weights each token using this exponential envelope with no extra forward passes, stabilized by difficulty damping and an entropy-gated fallback. CreditFilter outperforms GRPO and DAPO by 4–5 points on MATH500 and 8–10 on the AIME sets, and its dynamically adaptive variant matches it. The ablations reveal that the majority of this gain stems from ranking tokens within each response rather than globally across the batch, the coordinate on which the 80/20 mask ranks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.