acceptodds
Under review as a conference paper at ICLR 2027

When Is a Policy-Gradient Batch Reliable? Identification Limits Under Severe Uncertainty

Abstract

In reinforcement learning (RL), an agent learns a parameterized policy for sequential decision making by maximizing the expected cumulative return . Policy-gradient algorithms act on gradients estimated from finite trajectory batches, so a basic operational question is whether the resulting update points with high probability in the population-gradient direction. Existing work reduces gradient variance, adapts batch sizes, or controls policy improvement under explicit concentration assumptions, but these approaches do not answer a prior identification question: does the information conventionally retained about a one-trajectory gradient estimator determine its finite-batch wrong-direction probability? We show that, in general, it does not. Bounded atom-free REINFORCE problems can match any fixed jointly finitely generated menu of conventional one-rollout descriptors together with complete first- and second-order policy-gradient functions while having arbitrarily different strict directional risks. Under fixed Chernoff information and a bounded log-likelihood-ratio radius, we derive a sharp ambiguity law; a fixed pair of finite-horizon sequential REINFORCE Markov decision processes (MDPs) retains distinct directional large-deviation exponents despite exact matching of the stated policy-gradient and environment descriptors. We then give a non-oracle finite-sample certificate (without supplying the true population-gradient direction) using an independent confidence cone and held-out one-sided moment-generating-function information. Experiments show a regime-dependent picture: standard control settings can be close to Gaussian, whereas a controlled rare-critical family exhibits familywise-confirmed failures and a separately selected Grid2Op configuration survives an independent predeclared confirmation. The results position finite-batch policy-gradient reliability as an identification problem before it is a calibration problem.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.