acceptodds
Under review as a conference paper at ICLR 2027

Towards continual in-context learning via training cartridges at test time

Abstract

Language model agents can learn from experience through in-context learning (ICL), but ICL is limited by finite context window and declining performance as context grows. We formulate **continual in-context learning** as preserving ICL's performance and adaptability as experience grows beyond these limits. Our controlled experiments identify two requirements for iterative key-value (KV) cache compaction: each memory must preserve knowledge as new context arrives, and memories from successive compactions must work together. To meet these requirements, we introduce Accordion, which repeatedly consolidates an agent's oldest experiences into trained KV memory modules while keeping the base model frozen. The agent revisits these experiences to generate self-study examples, then distills them into a compact cache initialized with improved Attention Matching. Each new module is trained alongside frozen earlier modules and varying amounts of subsequent context, preparing it to remain useful as experience accumulates. Across QuALITY, Blind Spectrum Monitoring from Continual Learning Bench, and Harvey LAB, Accordion achieves relative improvements of approximately 13%, 60%, and 18%, respectively, over the strongest evaluated baselines using the same base model. It supports repeated compaction at - compression and handles tasks over corpora of 1M-12M tokens through compact memory and tool use.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.