acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Early-Turn Credit Failure in Agentic RL via Rollout-Free Reattribution

Abstract

Long-horizon capabilities are essential for LLM agents to solve complex tasks that require iterative tool use and multi-step interaction. Reinforcement learning, the predominant training paradigm for such agents, typically provides no intermediate rewards and reveals success only through a terminal reward. Learning from this sparse feedback requires turn-level credit assignment to estimate how each action contributes to the outcome and therefore determine its learning signal. In this work, we show theoretically and empirically that current critic-free methods suffer from **early-turn credit failure**, yielding unreliable credit in early turns. At these turns, Monte Carlo estimates are dominated by sampling noise, while final-outcome surrogate methods suffer from the weak connection between early actions and the outcome. Consequently, crucial early actions that shape the subsequent trajectory receive weak learning signals, limiting effective training on long-horizon tasks. To address this, we propose FOLIA, a rollout-free plug-in method that combines local influence with terminal hindsight for early-turn credit estimation. Specifically, FOLIA measures each action's local influence by counterfactually ablating the observed environment response, anchors this influence to the final outcome through a terminal hindsight score, and fuses the two signals with a confidence gate to reattribute credit across turns. Extensive experiments across diverse long-horizon agentic tasks show that FOLIA consistently improves existing RL algorithms through better credit assignment with negligible computational overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.