acceptodds
Under review as a conference paper at ICLR 2027

Decision-Certified Residual Sampling for Average-Reward Policy Verification

Abstract

An existing controller can perform well before the available observations justify accepting it. We study how historical data and additional samples can certify a policy in a finite average-reward Markov decision process. Two stationary-flow programs compare the best plausible gain with the worst plausible gain of the incumbent, giving an anytime-valid bound on policy loss. Separate dual potentials can improve this bound by an unbounded factor over the best common potential for the same incumbent. We then connect acquisition prices to actual sampling efficiency. With full-support dynamics, known rewards, and a unique optimal policy, normalized finite contractions consistently estimate the leading coefficients of the observed certificate excess above true policy loss, including its moving empirical centers. A price-tracking sampler realizes the optimal limiting allocation for this expansion and compares favorably with uniform sampling as the verification margin shrinks. A residual extension characterizes how heterogeneous historical counts reduce the additional samples. General validity and communicating-MDP completion require weaker conditions. Experiments distinguish certificate choice, allocation, policy quality, and computation, using controllers learned from behavior-policy logs and tests of the predicted stopping constants. An exact shared-model comparison also identifies uncertainty that the separate certificate can unnecessarily resolve.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.