Exploration and Exploitation Are Measurable for Language Model Agents
Abstract
Language Model (LM) agents are increasingly used in complex open-ended decision-making tasks, from AI coding to physical AI, where they must both explore the problem space and exploit acquired knowledge effectively. However, it remains challenging to systematically distinguish exploration and exploitation using observed actions without access to the agent's internal policy, and to determine when either constitutes an error. To address this, we design partially observable 2D grid maps paired with an unknown symbolic task directed acyclic graph (DAG), abstracting the structure of practical agentic AI scenarios. We then propose a metric that detects unproductive traversals using redundancy criteria grounded in classical graph exploration bounds, while permitting necessary backtracking. We attribute each error to exploration or exploitation based on whether the current state admits productive exploration, exploitation, or both, making the metric policy-agnostic. Using programmatically generated maps and task DAGs with controllable difficulty, we evaluate 35 frontier LM agents and find that even state-of-the-art models struggle to reach the goal within the step budget. Log exploration and exploitation errors exhibit strong linear relationships with task success ( and , respectively), providing empirical evidence that the metric captures behavior relevant to task success. Beyond success rate, the metric reveals distinct exploration and exploitation behaviors among models with similar success rates. We further find that a compact memory harness achieves higher task success than a reasoning baseline that retains the model's own reasoning traces. Finally, we extend our metric to an iconic video game, Minecraft Clone, beyond our synthetic environment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.