acceptodds
Under review as a conference paper at ICLR 2027

Almost Tight Regret Bounds for Reinforcement Learning with Abstention

Abstract

We study reinforcement learning (RL) with abstention, where an agent in a finite-horizon Markov decision process (MDP) may terminate an episode at any time and receive a fallback reward. We formulate this setting as an augmented MDP and propose an optimism-based algorithm that incorporates abstention into both value iteration and policy execution. We establish an interaction-dependent regret upper bound whose leading term depends on the realized number of interactions rather than the full decision budget , yielding a potentially sharper upper bound when abstention reduces interactions. To further characterize the MDPs in which such improvements occur, we introduce the optimal effective horizon length and the minimum sub-optimality gap over base actions. Our analysis shows that a shorter effective horizon can reduce regret, whereas a smaller gap increases the exploration cost. These effects are further illustrated through three concrete examples, including the boundary cases in which abstention is never beneficial or is uniformly optimal, as well as an intermediate setting in which the benefit of abstention varies with the problem parameters. Finally, complementary regret lower bounds show that dependence on and is unavoidable in general, supporting their role in characterizing the intrinsic learning difficulty of RL with abstention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.