acceptodds
Under review as a conference paper at ICLR 2027

Clipped Diagonal Natural Policy Gradient

Abstract

Natural policy gradient (NPG) methods account for the local geometry of the policy distribution, but computing directions preconditioned by the Fisher information matrix is expensive. Approximating this matrix with its diagonal yields a computationally efficient NPG algorithm, but such an algorithm has shown weak and sometimes unstable empirical performance in prior studies. We show that inverse-diagonal scaling can create nonlocal parameter displacements, while diagonal normalization can misestimate the update's length under the full Fisher metric. We introduce Clipped Diagonal NPG (CD-NPG). It clips each coordinate of the inverse-diagonal direction at a multiple of the median coordinate magnitude within its layer. It then uses the current minibatch to estimate the clipped direction's full-Fisher quadratic without forming the matrix. The method uses this estimate to rescale the direction to a target local policy change. These modifications preserve the asymptotic computational and memory complexities of diagonal NPG. Under matched hyperparameter-search budgets on six MuJoCo tasks, CD-NPG obtains returns comparable to or higher than policy gradients optimized with Adam or Muon on most tasks. Comparisons with two richer, more expensive NPG approximations are task- and training-setting-dependent. Diagnostics show substantially better local-KL calibration than diagonal normalization, while clipping ablations show that the clipped direction has fewer extreme finite-step departures from the local Fisher model than the unclipped direction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.