When Policy Improvement Breaks Future Learning: Self-Validating Policy Optimization
Abstract
Online reinforcement learning can invalidate its own next update. In continuing tasks with unbounded states, an approximate policy update can make the induced on-policy process heavy-tailed or poorly mixed. Subsequent value estimates and policy gradients may then become unreliable, even when the initial policy is stable and a stable optimum exists. To address this failure, we introduce a graded certificate for policy updates that identifies whether a raw policy preserves the statistical conditions needed by subsequent learning. We then develop self-validating policy optimization (SVPO), an average-cost actor–critic method that couples policy improvement to the validity of its own future updates. By retaining the physical task cost and preserving the certificate in the learned policy itself, SVPO keeps task performance aligned with the original objective and makes that policy valid for continued data collection and learning. The same construction yields Tilt, a certificate-preserving policy correction for existing policy optimizers. We prove that the certificate ensures that SVPO preserves the moment and mixing guarantees needed by future learning uniformly along its adaptive policy sequence. This establishes almost-sure convergence of the coupled actor–critic recursion to projected stationarity under standard assumptions. Experiments on switched LTI, input-saturated LQR, and unbounded queueing systems show that certificate-margin loss precedes downstream degradation, while SVPO maintains positive raw-policy margins with favorable physical task performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.