acceptodds
Under review as a conference paper at ICLR 2027

No Reset: Evaluating and Training Persistent Agents on Task Sequences

Abstract

AI agents are typically evaluated one task at a time, yet deployed agents operate in persistent sessions where conversation history, memory, files, and prior experience carry across tasks. We study this mismatch by introducing sequential variants of GAIA2 and SWE-Bench Pro, in which agents solve sequences of up to 40 tasks within a single persistent session. Our experiments show that persistence has a model-dependent effect on task success. Opus 4.6 and GPT-5.4 remain near their independent baselines, whereas DeepSeek v4 Pro loses 5.9 percentage points on average. More consistently, persistence changes how every model behaves: agents use up to 60% fewer output tokens, repeat less exploratory work, and, on GAIA2, increase Python's share of tool calls by up to . Trace analyses reveal both sides of this adaptation: agents reuse knowledge and customized tooling across tasks, but they can also under-explore new tasks and carry unchecked assumptions forward. To improve model capability in this setting, we introduce Sequential RL (Seq-RL), which trains on persistent multi-task rollouts with fine-grained credit assignment. Trained only on GAIA2-style data disjoint from the evaluation set, Seq-RL substantially narrows the sequential gaps on both GAIA2 splits. It also transfers zero-shot to SWE-Bench Pro, where it outperforms a single-task RL baseline trained on the same data, despite receiving no software-engineering data during RL. In short, sequential evaluation exposes new dimensions of agent reliability and in-context learning, while Seq-RL provides a promising training method for more robust long-lived agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.