Nonuniform Monte Carlo OPI Can Cycle
Abstract
Tsitsiklis (2002) proved convergence of discounted Monte Carlo optimistic policy iteration under uniform state selection and left convergence under fixed nonuniform state-selection frequencies open. We answer this question negatively for the scalar-stepsize, unnormalized asynchronous state-value recursion with model-based greedy improvement. For an explicit three-state, two-action MDP, nonuniform updates turn the greedy-policy mean dynamics into a rigorously certified stable six-policy cycle (a hybrid periodic orbit). With a bounded unbiased Monte Carlo return, we prove that, from a nonempty open set of initial conditions, sufficiently small square-summable Robbins–Monro stepsizes keep the original stochastic recursion near this cycle with any prescribed probability 1 − δ; the recursion therefore fails to converge on that event. The obstruction is geometric: uniform sampling produces radial residual contraction, whereas nonuniform scalar updates introduce an anisotropic distortion that can sustain policy switching. As a matched control, inverse-frequency normalization cancels this distortion and recovers a known almost-sure convergence guarantee for every strictly positive sampling law. The cycle persists on a relative open parameter neighborhood, the complete interval certificate succeeds independently at two nearby rational sampling laws, and continuation and paired stochastic experiments delineate the scalar-versus-normalized boundary.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.