acceptodds
Under review as a conference paper at ICLR 2027

Provably Efficient Ensemble Sampling for Reinforcement Learning

Abstract

Ensemble sampling offers a simple and practical approach to exploration in reinforcement learning (RL), yet establishing regret guarantees is challenging when persistent perturbations influence subsequent data collection. We propose an ensemble sampling framework with two variants that maintain persistently perturbed datasets, uniformly select one ensemble member per episode, and execute its greedy policy. By retaining historical perturbations and fitting only the selected member, the proposed variants avoid repeated noise resampling and iterative posterior sampling, offering a simple design amenable to deep RL. We establish regret guarantees for both variants in linear Markov decision processes (MDPs). The signed variant achieves a leading regret term of with ensemble members, matching the leading regret rate of a Thompson sampling-type method for linear MDPs while retaining perturbations across episodes. Here, is the feature dimension, the episode horizon, and the total number of interaction steps. The absolute variant improves this leading regret term to with only members, at the cost of an additional reference regression at each step. Our analysis handles the dependence between persistent perturbations and adaptively collected data by establishing stochastic optimism through joint perturbation control across steps for the signed variant and recursive ensemble-average optimism for the absolute variant. Experiments demonstrate that the proposed algorithms achieve performance competitive with existing baselines despite the simplicity of their ensemble sampling mechanism.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.