acceptodds
Under review as a conference paper at ICLR 2027

When History Fails to Become Experience: Action Calibration in Language Agents

Abstract

Language agents should draw on prior attempts and environmental feedback to improve subsequent decisions within the same task. However, providing additional interaction history can sometimes reduce task success, suggesting that agents do not consistently use this information effectively. To investigate this limitation, we conduct further experiments and find that, although history improves overall task completion, much of this benefit persists even when past actions are shuffled. Disrupting the correspondence between actions and observations causes only a modest decline in task success. We therefore hypothesize that agents do not reliably connect past actions with their outcomes when deciding how to proceed. To test this hypothesis, we explicitly label each returned observation as the outcome of the preceding action. This simple annotation improves task success and reduces next-action repetition without introducing new environmental information. Building on this insight, we introduce a learned calibrator that explicitly reassesses past experience and selectively records it to guide subsequent decisions, yielding further improvements in task success beyond outcome labeling alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.