acceptodds
Under review as a conference paper at ICLR 2027

Globally Convergent Offline Reinforcement Learning with Smoothed Bellman Residual Minimization

Abstract

Bellman Residual Minimization (BRM) enforces pointwise Bellman consistency and remains stable under neural parameterization, yet it is rarely used in offline reinforcement learning. Recent work shows that the debiased objective, though nonconvex minimax, admits a Polyak–Łojasiewicz–strongly-concave landscape (Kang et al., 2025), but only for offline inverse RL, and only up to the residual, which is not known to control value error. We propose Off-GLADIUS, which carries this geometry to offline RL, and establish global convergence for the procedure actually run. We then close the remaining gap by proving a local metric-subregularity inequality for the soft Bellman fixed point that converts residual excess risk into a finite-sample bound on value error under the optimal policy, assuming realizability and single-policy concentrability. Empirically, Off-GLADIUS is competitive with CQL and ahead of OptiDICE and prior BRM methods on three discrete-action control benchmarks. Carrying no pessimism term, density ratio, or target network, it also degrades far more gracefully as the logged data turns adversarial.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.