RACE: Recovery-Aware Credit Estimation for Long-Horizon LLM Agents
Abstract
Group-based reinforcement learning (RL) enables critic-free post-training of large language model (LLM) agents for long-horizon tasks. However, uniform trajectory-level credit conflates useful steps with unproductive actions, while GiGPO-style exact-state grouping misses outcome evidence from semantically comparable contexts with different observations. In this work, we introduce Recovery-Aware Credit Estimation (RACE), a plug-in method that assigns action-level credit through changes in contextual recoverability. Using a frozen text encoder, RACE nonparametrically estimates contextual recoverability from outcome-labeled neighboring contexts retrieved from previously collected rollouts, without training a critic or process reward model. A leave-one-trajectory-out retrieval scheme excludes the query trajectory from its retrieved neighborhood, and the clipped change in estimated recoverability across each transition serves as an action-level credit signal added to the base group-relative advantage. The method requires no additional environment interactions and supports both binary and graded outcomes. Experiments with Qwen3-4B and Qwen3-8B on ALFWorld, WebShop, and ScienceWorld show improvements over matched GiGPO runs across all three benchmarks. With Qwen3-8B, RACE achieves absolute success-rate gains of 7.85% and 7.59% on ALFWorld’s in-distribution and out-of-distribution splits, respectively, and improves the ScienceWorld graded score by 6.04 points and exact success rate by an absolute 14.75%. Further experiments show higher validation-curve AUC across all six benchmark–model settings and improved zero-shot transfer, measured by macro-averaged success and progress across BabyAI, PDDL, and TextCraft.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.