Finite Rollout Groups as Score-Conditioned Policy-Gradient Oracles
Abstract
A finite rollout group determines which population policy-gradient fields are available, not only how accurately they are estimated. We study iid on-policy groups over finite outcomes with outcome-measurable, permutation-equivariant, stop-gradient coefficients and no direct policy input. Conditioning leaves G − 1 peers: the conditional fields are exactly polynomials of degree at most G−1. A common-baseline quotient separates field representation from objective existence. Ordinary reward standardization has an exact boundary: for every finite G ≥ 2 and every positive strictly increasing variance denominator, a scalar simplex potential exists if and only if rewards have at most two distinct levels. Count-based rules yield exact self-inclusion and centering transformations. Under affine reuse of two fixed coefficient rules across linear reward weights, with positive response, we further prove Θ(G−2) minimax loss at the best surrogate optimum over a fixed strongly concave target class. Controlled Qwen2.5 studies at 1.5B and 7B identify count-rule objectives: all 65 objective-map endpoints are nearest their assigned optimum among three analytic candidates. Separate tests recover a rollout threshold, self-inclusion fixed points, and the predicted zero-to-nonzero circulation transition from two to three reward levels under ordinary standardization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.