When History Fails to Become Experience: Action Calibration in Language Agents
Abstract
Language agents should draw on prior attempts and environmental feedback to improve subsequent decisions within the same task. However, providing additional interaction history can sometimes reduce task success, suggesting that agents do not consistently use this information effectively. To investigate this limitation, we conduct further experiments and find that, although history improves overall task completion, much of this benefit persists even when past actions are shuffled. Disrupting the correspondence between actions and observations causes only a modest decline in task success. We therefore hypothesize that agents do not reliably connect past actions with their outcomes when deciding how to proceed. To test this hypothesis, we explicitly label each returned observation as the outcome of the preceding action. This simple annotation improves task success and reduces next-action repetition without introducing new environmental information. Building on this insight, we introduce a learned calibrator that explicitly reassesses past experience and selectively records it to guide subsequent decisions, yielding further improvements in task success beyond outcome labeling alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.