acceptodds
Under review as a conference paper at ICLR 2027

Preserving Pretrained Capabilities during RLHF via Cumulative Fisher Regularization

Abstract

Reinforcement learning from human feedback can improve preference satisfaction while degrading capabilities acquired before alignment. We introduce a Fisher-based update correction for proximal policy optimization (PPO) to address this capability-preservation problem. Our method estimates the initial policy’s output Fisher on capability-relevant prompts and corrects optimizer-proposed updates through a quadratic penalty on cumulative parameter displacement. Unlike a stepwise penalty, this formulation accounts for the interaction between each new update and the policy’s existing displacement from initialization. We derive a closed-form correction and a low-rank implementation that avoids constructing the full Fisher matrix. The correction selectively attenuates displacement along directions with high estimated output sensitivity, without requiring benchmark evaluation at every training step. We establish an evaluation protocol combining paired update diagnostics with PPO training trajectories, comparing against standard PPO, uniform update scaling, and activation-based correction. Mathematical task accuracy and frozen reward-model scores jointly assess capability retention and alignment gains, while held-out behavioral divergence tests the fidelity of the sensitivity estimate. This work provides a concrete optimization framework for preserving pre-existing capabilities during preference optimization, with cumulative behavioral sensitivity as its organizing principle.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.