acceptodds
Under review as a conference paper at ICLR 2027

Critic Error and Score Geometry in Token-Level Policy Gradients

Abstract

Token-level credit information need not reduce policy-gradient noise when token scores share parameters. We study this dependence in a deterministic, terminal-reward Token MDP. The advantage sequence is a martingale difference sequence, yielding an exact return-variance identity and a trajectory-gradient decomposition into an oracle token term and zero-mean cross-noise. The cross term has no fixed sign: a valid four-step shared-parameter policy reverses the second-moment ordering. Pathwise orthogonal scores permit an attainable ratio, or under homogeneous advantages when both estimators access the true value function. A spectral condition extends these conditional comparisons to random score geometry, but is uninformative on our measured neural-policy Grams. For estimated critics, a conditional mean-squared-error identity separates critic bias, its interaction with oracle error, and the full critic-error covariance; perturbation bounds and an exact-value GAE() spectrum complement this analysis. A pretrained GPT-2 prefix-continuation study exhibits a critic-budget-dependent estimator gap at substantial sampling cost. A separate three-seed, token-budget-matched LoRA diagnostic uses conditional- weights rather than oracle advantages and does not establish a training advantage. The results identify assumptions needed for an estimator separation and measurements needed before interpreting it as a practical efficiency gain.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.