Where Should the Credit Lie? Attention-Induced Anchor-Bridge Attribution Policy Optimization for LLM Reasoning
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of large language models. Methods such as GRPO provide critic-free optimization, but they assign uniform, sequence-level credit to all tokens and thus overlook the disparate roles that tokens play in the reasoning process. To address this, we leverage the model's own attention patterns to construct an attribution map that characterizes how information and contributions flow among tokens. This allows us to identify two pivotal token types: Anchor Tokens, which are frequently reused by subsequent reasoning steps, and Bridge Tokens, which connect preceding context with later reasoning. Building on these observations, we propose Anchor-Bridge Attribution Policy Optimization (ABPO), a token-level credit assignment framework guided by attribution relationships. Specifically, ABPO computes an AnchorScore and a BridgeScore for each token using the AnchorRank and BridgeTrace algorithms, respectively, and then fuses them into a SynergyScore for advantage reweighting. Experimental results across diverse reasoning tasks demonstrate that ABPO significantly and consistently outperforms the GRPO baseline and other credit assignment methods (+2.6 points on math, +7.0 points on other domains vs. GRPO). Furthermore, ABPO generates more concise reasoning chains (-3.15% tokens on math, -30.15% on other domains over GRPO), and achieves faster training convergence, highlighting its dual strengths in both effectiveness and efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.