Ensemble Sampling with Optimistic Refreshes for Behavior Foundation Models
Abstract
Online task inference with behavior foundation models (BFMs) aims to retrieve a high-performing policy for a target reward function through interaction with the environment, without updating the pretrained model. However, collecting reward-labeled data while achieving high returns at low computational cost remains challenging. In this paper, we propose Ensemble Sampling with Optimistic Refreshes for BFMs (ESOR-BFM), an online task inference method that combines ensemble sampling with optimistic recentering. When the regularized Gram matrix built from state features changes substantially, ESOR-BFM refreshes the ensemble perturbations and recenters the ensemble around an optimistic task embedding estimate. ESOR-BFM samples ensemble members for policy selection between refreshes and performs optimistic search only at refresh times, amortizing the search cost over multiple interactions. We theoretically establish a sublinear regret guarantee for an episodic variant of ESOR-BFM with well-trained BFMs, matching the state-of-the-art regret rate for ensemble sampling in stochastic linear contextual bandits. Extensive experiments on ExORL show that ESOR-BFM consistently outperforms existing online task inference methods and achieves the highest mean returns across settings with stationary rewards, selective reward labeling, and inaccurate successor feature estimates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.