Adaptive Behavioral Foundation Models: Continual Training from Reward-Free Experience
Abstract
Behavioral foundation models (BFMs) learn general-purpose policies and successor features from reward-free offline data, enabling agents to retrieve a policy for a given reward function without additional learning or planning. However, this policy can fail when test-time dynamics differ from those seen during pretraining. Memory-based approaches address this mismatch through in-context learning without updating model parameters. We study reward-free continual training for adaptation to unseen dynamics within a limited online interaction budget. We introduce AdaBFM, which frames pretraining as learning an adaptive prior distribution over test-time dynamics. At test time, it adapts this prior through continual fine-tuning on reward-free online experience. The zero-shot policy is conditioned on causal memory and a belief over successor ensemble members, allowing it to adapt as evidence accumulates, while successor-belief information gain guides continual training. We evaluate AdaBFM across an aggregate of 23 navigation, locomotion, and manipulation tasks under unseen dynamics. AdaBFM achieves the highest mean performance in 9 out of 23 tasks in the zero-shot setting and 16 out of 23 after continual training, improving by an average of 28% on harder tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.