Preserving Pretrained Capabilities during RLHF via Cumulative Fisher Regularization
Abstract
Reinforcement learning from human feedback can improve preference satisfaction while degrading capabilities acquired before alignment. We introduce a Fisher-based update correction for proximal policy optimization (PPO) to address this capability-preservation problem. Our method estimates the initial policy’s output Fisher on capability-relevant prompts and corrects optimizer-proposed updates through a quadratic penalty on cumulative parameter displacement. Unlike a stepwise penalty, this formulation accounts for the interaction between each new update and the policy’s existing displacement from initialization. We derive a closed-form correction and a low-rank implementation that avoids constructing the full Fisher matrix. The correction selectively attenuates displacement along directions with high estimated output sensitivity, without requiring benchmark evaluation at every training step. We establish an evaluation protocol combining paired update diagnostics with PPO training trajectories, comparing against standard PPO, uniform update scaling, and activation-based correction. Mathematical task accuracy and frozen reward-model scores jointly assess capability retention and alignment gains, while held-out behavioral divergence tests the fidelity of the sensitivity estimate. This work provides a concrete optimization framework for preserving pre-existing capabilities during preference optimization, with cumulative behavioral sensitivity as its organizing principle.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.