acceptodds
Under review as a conference paper at ICLR 2027

Quantile of Means: A Bonus-Free Ensemble Method for Minimax Optimal Reinforcement Learning

Abstract

Optimal Reinforcement Learning (RL) algorithms typically rely on carefully constructed count-based uncertainty estimates, such as bonuses or posterior distributions, to drive exploration. Although theoretically sound, such estimates are difficult to scale and therefore offer limited insight for designing exploration heuristics in practice. Meanwhile, ensembling has emerged as a practical approach, but remains without theoretical justification. Building on a recent ensemble-based method for Multi-Armed Bandits, we propose a quantile-based ensemble method for finite-horizon tabular MDPs. Our simple bonus-free approach achieves optimal variance-dependent regret bounds without explicit uncertainty estimation, providing theoretical grounding for ensemble-based exploration in RL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.