Direction-Preserving Credit Redistribution for Long-Horizon LLM Agents
Abstract
Long-horizon LLM agents must infer which individual actions deserve credit from sparse trajectory-level outcomes. Existing step-level estimators provide finer-grained supervision, yet the resulting credit is not uniformly reliable: which parts of these signals should the policy trust? We reveal a complementary asymmetry between non-parametric and parametric credit. Non-parametric advantages, derived from realized-return contrasts, provide informative update directions, but their magnitudes remain coupled to sampled continuations and can misallocate credit across actions. Conversely, accurate value or reward estimation is difficult in long-horizon agentic tasks, causing parametric prediction errors to propagate into policy advantages and even reverse update directions; nevertheless, learned value movement retains useful information about relative action contribution. Based on this observation, we introduce Grounded Relative-Value Advantage Weighting (GRAW), a plug-in for non-parametric step-level estimators. GRAW retains the base credit as a grounded allocation prior and uses learned relative value movement only to reweight actions within the same update direction, preserving every base advantage sign and the total absolute credit in each direction. Across three long-horizon agent benchmarks and multiple model scales, GRAW consistently improves strong non-parametric step-level baselines, achieves the best overall performance, reduces variance across seeds, and adds only minimal computational overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.