Token Credit Assignment as a Boundary-Value Problem in RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) provides sparse outcome-level feedback, while methods such as GRPO assign the same credit to every token in a response. With binary terminal rewards, exact local hindsight evidence and the initial value determine a Bellman value path whose increments are token advantages and whose endpoints fix their total budget. Motivated by on-policy distillation (OPD), teacher–student likelihood ratios provide surrogate evidence, but their bias and noise can cause the reconstructed path to violate this budget or value boundary. To reconstruct the Bellman value path guided by OPD, we propose three design requirements targeting the objective of RLVR, utilizing the information of OPD. These requirements lead to Endpoint-Locked Value-path Evidence Reconciliation (ELVER). ELVER-Affine projects raw credits onto the budget hyperplane in closed form, while ELVER-Box reconstructs a path with constrained value bounds, and it minimizes squared Euclidean distances to the evidence-induced transition lines, yielding a strictly convex quadratic program. Experiments on five mathematical reasoning benchmarks show that ELVER-Box outperforms standard methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.