CFPP: Conflict-Free Preservation Projection for Locate-then-Edit Model Editing
Abstract
Model editing provides an efficient way to update factual knowledge in large language models (LLMs) after deployment, without retraining the full model. A widely used locate-then-edit paradigm localizes critical MLP representations and computes a targeted parameter update. However, its preservation term can penalize edit-relevant directions when the dominant subspace of its covariance overlaps with the active edit subspace. This preservation–edit conflict becomes more severe as the number of edits grows, leading to under-editing and degraded generalization. We propose CFPP (Conflict-Free Preservation Projection), a lightweight framework that modifies the preservation covariance before computing the local closed-form update. CFPP stabilizes the edit geometry through spectral whitening and applies residual-aware soft projection with trace calibration to attenuate edit-aligned components while maintaining the overall preservation budget. CFPP is designed for locate-then-edit methods such as MEMIT, PMET, and EAMET. Experiments on GPT2-XL, Mistral-7B, and LLaMA3-8B show that CFPP yields an average relative gain of 8.2% in the reported Avg. score, with improvements reaching 31.5% in individual settings. Code is available at https://anonymous.4open.science/r/CFPP-B4H7.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.