Average-Reward Policy Optimization via Discounting: Anytime and Global Linear Convergence
Abstract
We study average-reward policy optimization via discounting in finite Markov decision processes where every stationary policy induces one recurrent class, possibly with transient states. With a known model and exact policy evaluation, Discount-Guided Policy Gradient achieves an anytime optimality gap after policy updates, with nondecreasing average rewards, from any initialization and for any fixed positive discounted step size. The proof combines local gradient domination with an explicit bound on entry into a near-optimal region. For rational model data, Blackwell Policy Mirror Ascent achieves global R-linear convergence of the average-reward gap from every full-support initialization through a multiplicative error bound at an explicit Blackwell discount. Neither guarantee requires a finite global stationary-distribution mismatch coefficient.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.