Beyond Memory: Disentangling the Persistence of Task Gaming in AI Agents
Abstract
Agentic systems can sometimes exploit weaknesses in their evaluators rather than satisfying the intended task. This behaviour, where an agent satisfies the evaluator while departing from the intended task, is known as task gaming. But what makes such behaviour persist or change when an agent continues working? We study this by varying what an agent retains when resuming a task, including its conversation history, saved work and task-relevant information. In controlled paired experiments, restart context changes constraint violations even when agents begin from identical, correct code. Retaining prior work also reduces the effort required to continue, although control experiments show that this reuse is not specific to gaming. In a separate task, providing information acquired during an earlier attempt enables successful completion without further detected shortcuts, even with minimal conversation history. However, explicitly explaining changes made to the restored workspace does not reliably reduce violations, so the mechanism behind the context effect remains unclear. Finally, our audit results reveal that apparently correct solutions can manipulate the evaluator through surrounding code, exposing violations that checks of the intended function alone would miss. Altogether, these results separate three effects of retained experience: changes in constraint adherence, reuse of prior work and access to useful information. They show that what an agent retains from earlier attempts can meaningfully shape its subsequent behaviour.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.