acceptodds
Under review as a conference paper at ICLR 2027

Focal Credit Assignment for Ultra Long-Horizon Agentic Reinforcement Learning

Abstract

Training long-horizon large language model (LLM) agents with outcome-based reinforcement learning poses a credit-assignment challenge. GRPO assigns the same trajectory-level advantage to all actions, despite their different contributions to task success. Through a policy-gradient analysis, we identify two sources of mismatch between this trajectory-level advantage and the ideal action-specific credit: value shifts inherited from preceding history and stochasticity in subsequent interactions. Empirical analysis further reveals that sharp value changes are concentrated at a small fraction of actions. Motivated by these observations, we propose FocalRL, which tackles long-horizon credit assignment through local action repair. Specifically, an LLM-based locator first identifies consequential errors in failed trajectories and generates rubrics specifying the desired corrections. Then, we restore the interaction state before each identified error and sample bounded local continuations, which are evaluated against the repair rubrics. We combine these rubric-scored local groups with outcome-scored full-trajectory groups for policy optimization, providing targeted supervision that mitigates history value offsets and reduces the influence of downstream stochasticity. Experimental results show that FocalRL outperforms GRPO and most related methods under same rollout budget on representative agentic search and coding benchmarks, including BrowseComp, GAIA, and SWE-bench Verified. Notably, our deep research agent built on Qwen3.5-9B achieves 57.9% accuracy on BrowseComp, setting a new state-of-the-art among sub-10B models. Our code is available at https://anonymous.4open.science/r/FocalRL-5DE0/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.