acceptodds
Under review as a conference paper at ICLR 2027

Mixing Regularization for Average-Reward -Learning and -Planning

Abstract

The average-reward setting for reinforcement learning is challenging because its Bellman operator is known to contract in span only when the underlying Markov chain is mixing. Whereas existing work typically a contraction property, e.g. through environment mixing, we consider the setting where we are given a reward-agnostic -step mixing policy, and we the contraction and exploration properties of the learned policy by regularizing toward the mixing policy. We provide finite-time analysis of -learning algorithms that can be instantiated with multiple choices of regularization, with guarantees measured with respect to the regularized fixed point. Our framework exposes novel regularizer-dependent objects that control contraction, exploration, and operator bias. In the one-step setting, our rates match the best known rates for -iteration and synchronous -learning, and for -step mixing we derive new rates. We also study -planning variants, and show that planning acts as time-localized learning, where the planning horizon trades off the bias from the boundary value against the sample complexity. Finally, we present numerical studies validating the analysis: (i) knowledge of a strong mixing policy improves convergence and exploration, and its advantage increases with bottleneck severity and (ii) our theory predicts the convergence and bias rankings of regularizers from the -divergence family.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.