acceptodds
Under review as a conference paper at ICLR 2027

Fast Rates for Offline Q-Learning with KL Regularization

Abstract

This paper investigates offline Q-learning for finite-horizon Markov decision processes with Kullback-Leibler (KL) regularization. Recent studies have shown that KL regularization can yield fast rates in offline learning. However, it remains unclear whether Q-learning, a stochastic approximation style algorithm, can achieve such rates in the offline setting. We first propose KL-regularized Q-Learning (KLQ), a pessimism-free algorithm for offline KL-regularized reinforcement learning. By exploiting the quadratic relationship between policy suboptimality and Q-estimation error induced by KL regularization, we show that this quadratic error structure can be preserved through the Q-learning update. Under single-policy concentrability, KLQ achieves a sample complexity of to find an -optimal policy. This sample complexity, however, has an exponential dependence on the regularization parameter and the horizon. To remove this explicit exponential dependence, we further propose KL-regularized Pessimistic Q-Learning (KLPQ), which incorporates a pessimistic bonus and a monotone Q-update. We prove that KLPQ retains the same sample complexity while removing this exponential dependence. To the best of our knowledge, these results provide the first finite-sample guarantees for Q-learning that establish a \(1/\epsilon\) dependence on the target accuracy for offline learning in tabular finite-horizon KL-regularized MDPs under single-policy concentrability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.