acceptodds
Under review as a conference paper at ICLR 2027

Average-Reward Policy Optimization via Discounting: Anytime and Global Linear Convergence

Abstract

We study average-reward policy optimization via discounting in finite Markov decision processes where every stationary policy induces one recurrent class, possibly with transient states. With a known model and exact policy evaluation, Discount-Guided Policy Gradient achieves an anytime optimality gap after policy updates, with nondecreasing average rewards, from any initialization and for any fixed positive discounted step size. The proof combines local gradient domination with an explicit bound on entry into a near-optimal region. For rational model data, Blackwell Policy Mirror Ascent achieves global R-linear convergence of the average-reward gap from every full-support initialization through a multiplicative error bound at an explicit Blackwell discount. Neither guarantee requires a finite global stationary-distribution mismatch coefficient.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.