Preference-Aware Reward Redistribution for Token-Level RLHF
Abstract
Outcome reward models (ORMs) provide robust preference signals for reinforcement learning from human feedback (RLHF), yet their scalar response-level rewards leave the contribution of individual token decisions unspecified, creating a granularity mismatch with token-level policy optimization. We observe that the ORM itself provides a natural source of weak supervision for resolving the resulting ambiguity in mapping response-level rewards to individual tokens: the outcome reward is predicted from its terminal hidden representation, while intermediate token representations exhibit different degrees of alignment with this terminal state. We exploit this representational relation as a proxy for token-level preference relevance and propose Single-Source Entropic Optimal Transport (SEOT), a training-free framework that formulates outcome-to-token reward redistribution as a globally normalized allocation problem. SEOT models the terminal hidden state as a Dirac source and response-token hidden states as target supports, with their representational similarity defining preference-aware transportation costs. The resulting single-source entropic formulation admits a closed-form solution and requires neither additional reward model training nor iterative OT optimization, while exactly preserving the original outcome reward. Experiments on mathematical reasoning and general-domain tasks demonstrate consistent improvements over existing reward redistribution methods. Further analysis shows that SEOT better emphasizes preference-critical tokens while avoiding overly concentrated reward allocation, providing more effective token-level guidance for policy optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.