AttnSR: Attention Step Reward for Long-Horizon Agentic RL
Abstract
Long-horizon LLM agents make multi-step decisions, yet terminal-reward-only reinforcement learning provides little direct supervision for individual steps. We introduce AttnSR, a step-level reward augmentation method that derives reward allocations from policy attention extracted during training replay, without a separately trained large reward model or additional environment rollouts. AttnSR traces backward through the policy model attention graph to estimate the relevance of earlier steps to the final response. The terminal step, measured relative to a running baseline, determines whether these rewards reinforce or penalize the selected steps. In a controlled Qwen2.5-7B-Instruct study, AttnSR consistently outperforms baseline across all evaluated checkpoints on ALFWorld, raising final held-out success from 91.2% to 97.1%. The step-reward ablation experiment shows that the benefit depends on how intermediate rewards are assigned, rather than simply on their presence. AttnSR also improves ScienceWorld at every checkpoint while cutting interaction by 29.9%. No statistically detectable performance improvement is observed on WebShop. These results support policy attention as a practical prior for step-level reward allocation, with benefits that may depend on task complexity and horizon.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.