Fisher-Feature Proximal Policy Optimization
Abstract
Proximal Policy Optimization (PPO) loses plasticity in its trunk, and the usual diagnostic is the rank of the penultimate feature matrix. Proximal Feature Optimization (PFO) follows that diagnostic by penalizing feature drift equally in every direction. We argue that the diagnostic is the weak link. The policy sees the features only through two linear heads whose width is set by the action space, so rank, computed without the heads, cannot separate representations the policy treats very differently; and rank is a count settled at the tail of the spectrum, so it misses how variance is shared. We derive how an update on one state leaks into the clipped ratio of another: the leakage is governed by the cross-state coupling of the features, their squared inner products, which depends on how the spectrum is spread and on the feature norm, and not on a count. We measure the angular part by the participation ratio of the projected outputs, essentially the inverse mean squared cosine between states, and replace the isotropic penalty by the unique quadratic penalty that measures policy change, the Fisher information pulled back through the heads (FF-PPO). On six MuJoCo tasks the predicted mechanism is visible directly: FF-PPO lowers the coupling by more than an order of magnitude relative to PPO and several-fold relative to PFO, and wherever the penalties help, the methods order by coupling exactly as they order by trust-region excess. The gain in return grows with the width of the policy-visible subspace, where FF-PPO separates from PFO, and with the strength of the non-stationarity: with more optimization epochs per batch or under changing dynamics, PPO collapses while FF-PPO keeps learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.