acceptodds
Under review as a conference paper at ICLR 2027

Focus on the Critical: Multi-Granularity Modulated Credit Assignment for Long-Horizon LLM Agents

Abstract

Efficient credit assignment is critical for reinforcement learning (RL) of large language model (LLM) agents in long-horizon interactive environments, where supervision is typically sparse and delayed. However, group-based policy optimization methods often apply a trajectory-level signal uniformly across steps and tokens, which can obscure a small number of pivotal decisions that determine task success. To address this limitation, we propose a *Multi-Granularity Modulated* (MGM) credit assignment framework. It refines the learning signal by integrating token-level Top- entropy filtering to isolate critical decisions, step-level temporal modulation to differentiate early exploration from late-stage execution, and task-level weighting to prioritize challenging queries. Through this unified hierarchical design, MGM effectively highlights pivotal actions, enabling more accurate and efficient policy optimization. Extensive experiments on ALFWorld, WebShop, and Search-Augmented QA demonstrate that our method consistently outperforms state-of-the-art baselines, up to 25.0% absolute improvement based on GRPO and 14.5% based on DAPO. Moreover, we observe a robust unimodal sensitivity to the Top- ratio, with best performance typically achieved when retaining roughly –% of tokens for entropy estimation, suggesting that excluding a small fraction of low-entropy tokens can improve the signal-to-noise ratio in long-horizon credit assignment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.