acceptodds
Under review as a conference paper at ICLR 2027

Randomized Learning-Rate -Learning for Continuing Discounted MDPs

Abstract

Randomized learning-rate -learning achieves tractable, gap-independent regret in episodic finite-horizon MDPs and often outperforms optimistic-bonus -learning, but its guarantees rely on a terminal boundary that ends the Bellman recursion. Removing this boundary leads to continuing discounted control, for which no such guarantee is known. We propose a discounted randomized -learning algorithm (D-RandQ), whose analysis lifts stage-wise reuse counting to a global interaction-time occurrence representation: each target is aligned with its successor occurrence, stage-ending self-loops and boundary visits are charged separately, and the recursion closes through a reuse operator with bounded column sum. Combined with adaptive-update concentration and Dirichlet anti-concentration, this yields a finite-sample, high-probability, gap-independent action-gap bound, hence, to our knowledge, the first finite-sample, gap-independent sublinear regret guarantee for randomized learning-rate -learning in continuing discounted MDPs. Empirically, under a documented disjoint calibration protocol, D-RandQ attains lower cumulative action gaps and higher realized rewards than -SlowSwitch-Adv and UCBVI- on Maze-25, and a large separation on the reported Chain-50 runs at the same transferred operating point.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.