Rethinking Agentic Continual Learning as Bayesian Discovery
Abstract
Large language models (LLMs) are increasingly used as agents, yet they rarely turn test-time experience into persistent improvements. We connect agentic continual learning (CL) to Bayesian discovery: both accumulate evidence from sequential interactions into persistent knowledge. With the model frozen, what an agent learns can persist only as text, and we represent it as falsifiable hypotheses about the entities of its environment, whose beliefs each task updates by Bayes’ rule. This view clarifies what a task stream must provide: a refutation corrects rather than interferes, forgetting arises when memory revision discards evidence, and a stream becomes silent when later tasks provide no evidence for or against what was learned earlier. We propose DiscoverCL, a training-free harness that approximates these updates with a research loop of proposing hypotheses, acting, and recording evidence. We further introduce DiscoverScore, which measures how much a stream updates an agent’s beliefs per task. With Claude Sonnet 5 and Opus 5 on six CL-Bench streams, DiscoverCL matches or improves over the no-memory baseline in all 12 settings and is the strongest bounded-memory method in 9 of 12. With Opus 5, it recovers 50–97% of full-context ICL’s gain on Database, Spectrum, Sales, and Codebase while using 2.6–9.6× less memory. DiscoverScore ranges from 0.37 to 3.48 bits per task, showing that the streams differ in how much they allow the agent to learn continually.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.