acceptodds
Under review as a conference paper at ICLR 2027

Towards Regret for Private RL with General Function Approximation

Abstract

Recent work on private reinforcement learning with general function approximation establishes regret over episodes, but whether privacy permits the canonical dependence remains open. We resolve this question for a finite Bellman-complete value-function class with bounded Bellman–Eluder dimension. Our main algorithm, Priv-LazyGOLF, privatizes rare policy switching through a twofold use of AboveThreshold: it privately detects when the current hypothesis has accumulated enough Bellman inconsistency and privately selects an optimistic feasible replacement. We show that Priv-LazyGOLF is -jointly differentially private and, with high probability, achieves regret with only logarithmically many policy switches in . Thus, privacy preserves the regret of non-private learning. We also establish that the same private rare-switching template extends, through suitable decomposable losses, to a broader class of reinforcement learning problems while retaining regret and privacy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.