acceptodds
Under review as a conference paper at ICLR 2027

Fast and Slow Online Learning for Agents

Abstract

Large language model (LLM) agents are typically deployed as static models, yet their deployment conditions can differ from training and continue to evolve over time, requiring agents to continually adapt after deployment. Unlike offline post-training, however, every intermediate policy during deployment serves subsequent queries, making test-time learning essentially an online learning problem, where the learner is evaluated by its cumulative performance over the interaction stream rather than only its final performance. This cumulative objective requires both fast adaptation and strong long-run improvement. We therefore introduce Fast and Slow Online Learning (FSOL), a framework that combines memory for fast adaptation with RL for progressively internalizing experience into model weights. We evaluate FSOL on writing-style adaptation, search-augmented QA, WebShop, and ALFWorld across stationary streams, distribution shifts, mixed tasks, and sequential tasks. Across these settings, FSOL consistently achieves higher cumulative performance than memory or RL alone, rapidly adapts to distribution shifts, effectively learns from heterogeneous interleaved tasks, and retains previously acquired capabilities while adapting to new ones.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.