acceptodds
Under review as a conference paper at ICLR 2027

Continual Reinforcement Learning via Behavioral Foundation Models

Abstract

Continual reinforcement learning (CRL) requires agents that adapt quickly to new tasks, reuse past knowledge, and avoid forgetting. Existing methods address parts of this trade-off but rely on restrictive assumptions, such as slow gradient-based adaptation or known task boundaries. We ask whether Behavioral Foundation Models (BFMs), a recent class of zero-shot RL methods, can serve as a backbone for CRL. BFMs are appealing: switching to a new task-conditioned policy does not require task-specific gradient updates, the unsupervised objective is reward-agnostic, and the induced policy family may mitigate forgetting if coverage is maintained. However, BFMs are designed for offline settings and fail when applied naively online: zero-shot inference assumes a reward model, and online updates on a single task cause coverage collapse. To enable BFMs as continual online agents, we introduce a generic three-phase pipeline: reward-free offline pre-training; reward-model-free online task inference, which switches policies without updating their parameters; and continual updates through a coverage-preserving buffer. We evaluate continual sequences from DM Control and Continual World to characterize where BFMs succeed, where pre-training coverage becomes a bottleneck, and how each component contributes to stability and adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.