Deep Exploration as Sequential Testing in -Function Space
Abstract
Best-policy identification in Markov decision processes (MDPs) is fundamentally an information-acquisition problem: the learner must gather trajectories that resolve uncertainty about which sequences of actions are optimal. We formulate this objective as a sequential testing problem in the space of optimal -functions, where each policy induces a region of optimality and exploration seeks to identify the region containing the unknown Bellman fixed point. We introduce TQS, a model-free algorithm that implements this testing problem through a leader–challenger mechanism. At each round, TQS pairs an empirical estimate of with a challenger sampled from a Gibbs distribution based on empirical Bellman residuals, and uses the challenger's squared fitted Bellman residual as an adaptive exploration reward. Since this reward changes across rounds, TQS uses an online RL learner as a dedicated exploration controller, steering the agent toward trajectories that maximally separate the competing hypotheses. In linear MDPs, we show that TQS admits a saddle-point interpretation in which the behavior policy maximizes distinguishability against a randomized alternative, and we derive Gibbs-contraction and sample-complexity guarantees governed by an instance-dependent oracle rate. Unlike model-posterior methods, TQS operates directly in value-function space and extends naturally to neural function approximation. We instantiate a value-disagreement surrogate with Bayesian heuristics, and obtain strong results on large tabular and continuous state benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.