acceptodds
Under review as a conference paper at ICLR 2027

Protect What Predictions Rely On: Decision-Aware Preservation for Massive Model Editing

Abstract

Editing thousands of facts in a language model requires deciding which aspects of its existing behavior to protect. Closed-form editors commonly make this decision using background activation statistics, which measure representation usage but do not directly measure predictive sensitivity. We propose Decision-Aware Preservation (DAP), a reusable preservation metric that augments activation statistics with prediction-gradient information. DAP weights Wikipedia probe keys by squared loss-gradient norms and adds the resulting matrix to the host editor's preservation term. This construction follows from an upper bound on loss derivatives at selected probe positions and retains the host's closed-form solution without changing target optimization. Experiments on GPT2-XL, GPT-J (6B), and Qwen2.5-7B across two benchmarks show improved locality at comparable edit success with up to 10,000 edits. The evaluated compositions gain up to 9.4 percentage points in locality success and 4.3 points in overall editing score over their hosts. Matched-probe controls on Qwen/PRUNE and GPT-J/RECT at 10,000 edits show that correct gradient–key pairing improves preservation beyond uniform or shuffled weights, with gains persisting after independent strength tuning. Fixed-spectrum interventions further characterize the role of protection orientation. Together, these results identify sensitivity-guided allocation as a practical way to improve preservation in closed-form model editing. Code is available at https://anonymous.4open.science/r/dap-editing-46F4/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.