acceptodds
Under review as a conference paper at ICLR 2027

Continual Learning Exam

Abstract

Large language model agents are increasingly designed for continual learning beyond parameter updates, acquiring and updating knowledge from sequential experience while retaining useful prior knowledge. Current evaluations offer limited insight into this capability, as they largely assess individual task performance rather than improvement through accumulated experience. We introduce Continual Learning Exam (CLE), a benchmark developed by 50+ domain experts for measuring experience-driven improvement in realistic professional workflows. CLE comprises 120 human-authored tasks, each organized as a five-level ladder, for a total of 600 levels spanning 89 professional roles across 12 occupational domains. Agents must pass each level to unlock the next, with later levels building on experience accumulated during earlier assignments. Each submission is assessed by an evaluation agent against human-authored rubrics. We compare a single submission per level with a setting that also provides diagnostic feedback on unmet requirements after failure, allowing the agent to reflect on the feedback and resubmit once. Across 17 harness-model configurations, this feedback and revision opportunity yields a median completion gain of 5.3 percentage points. Yet even the best-performing system clears only 47.7% of levels and completes all five levels in just 23.3% of tasks, highlighting continual learning as a key challenge for current LLM agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.