acceptodds
Under review as a conference paper at ICLR 2027

Nonuniform Monte Carlo OPI Can Cycle

Abstract

Tsitsiklis (2002) proved convergence of discounted Monte Carlo optimistic policy iteration under uniform state selection and left convergence under fixed nonuniform state-selection frequencies open. We answer this question negatively for the scalar-stepsize, unnormalized asynchronous state-value recursion with model-based greedy improvement. For an explicit three-state, two-action MDP, nonuniform updates turn the greedy-policy mean dynamics into a rigorously certified stable six-policy cycle (a hybrid periodic orbit). With a bounded unbiased Monte Carlo return, we prove that, from a nonempty open set of initial conditions, sufficiently small square-summable Robbins–Monro stepsizes keep the original stochastic recursion near this cycle with any prescribed probability 1 − δ; the recursion therefore fails to converge on that event. The obstruction is geometric: uniform sampling produces radial residual contraction, whereas nonuniform scalar updates introduce an anisotropic distortion that can sustain policy switching. As a matched control, inverse-frequency normalization cancels this distortion and recovers a known almost-sure convergence guarantee for every strictly positive sampling law. The cycle persists on a relative open parameter neighborhood, the complete interval certificate succeeds independently at two nearby rational sampling laws, and continuation and paired stochastic experiments delineate the scalar-versus-normalized boundary.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.