acceptodds
Under review as a conference paper at ICLR 2027

Direction-Aware Token Credit Assignment with Reward-Model Gradients

Abstract

GRPO shares a response-level advantage across tokens, leaving unresolved how strongly each token should contribute to learning. Updates induced by the shared advantage reinforce or suppress the sampled token at each position, but these updates do not always point toward increasing reward at the token level. This motivates a **work-based view of credit**: just as work depends on the component of force along a displacement, token credit should reflect the local reward change along the policy update direction. We introduce **Direction-Aware Token Credit Assignment (DATC)**, which maps frozen reward-model gradients into policy logit space and allocates weights by evaluating the alignment between the reward direction and GRPO's signed sampled-action update direction, while preserving the sign and mean of the response-level advantage. DATC outperforms GRPO in all four settings across two model families. Our code will be released at https://anonymous.4open.science/r/DATC-0161.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.