acceptodds
Under review as a conference paper at ICLR 2027

No Do-Overs: Observation-Grounded Credit for Language-Model Agents Without Re-Execution

Abstract

Improving a long-horizon language-model (LM) agent is a problem of credit assignment: localizing a terminal outcome to the steps that earned it. We study it in the regime a deployed agent faces (credit assignment in deployed agents), where a text-action LM agent must learn from what it logged. A deployed agent updates its policy off-peak, while it is not serving user traffic, on the trajectories that serving already produced, and by update time cannot rewind a session, re-execute a tool call, or consult an oracle: its actions are irreversible. Yet the dominant tools for per-step credit in agentic RL rely on exactly the counterfactual access a deployed agent lacks, so credit must come from the logged trajectory alone. One signal the rule leaves readable is the observation itself, which in a text-action setting states in words what each action did. Prior methods consume observations too, but only through a learned model or a monolithic self-belief; we instead read credit directly off the observation, with no trained reward or value model on the credit path. We call this observation-grounded step credit. As a GRPO reward it improves long-horizon agents across three environments using only logged information. On ALFWorld, two rounds raise the base from to , the highest among same-backbone methods and on par with a privileged comparator that uses strictly more information; the same operator improves over its terminal-credit ablation and strong baselines on ScienceWorld.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.