acceptodds
Under review as a conference paper at ICLR 2027

Fisher-Trace State Sensitivity for Proximal Policy Optimization

Abstract

Deep reinforcement learning with neural critics often suffers from unstable value estimates and overly aggressive policy updates. Motivated by the intuition that states where the critic is highly sensitive deserve more cautious updates, we derive a state-level sensitivity measure for Proximal Policy Optimization (PPO) from the Fisher-information trace of the critic. We prove that, under the training conditions of standard PPO, this trace reduces to the squared norm of the value gradient with respect to states, and batch-wise standardization removes the unknown global scale. The resulting algorithm, GFIM-PPO, rescales the PPO surrogate loss with a bounded coefficient computed from this measure. On a diverse suite of reinforcement learning benchmarks, GFIM-PPO improves over PPO on the majority of tasks. Our analysis further confirms that it is the state-dependent construction of the weighting coefficient, rather than weighting alone, that drives the gains. Code is available at https://anonymous.4open.science/r/CAAC-F11F.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.