acceptodds
Under review as a conference paper at ICLR 2027

Revisiting the Dynamics of Target Optimization in Large Language Model Editing

Abstract

Knowledge editing has emerged as a promising approach for efficiently updating knowledge in large language models (LLMs). These methods first optimize a target hidden state that steers the model toward a new output, and then edit the model parameters to map the original hidden state to this target. While most existing methods focus on the second stage of parameter editing to preserve other knowledge, we find that the geometry of upstream target hidden-state optimization is also critical to editing locality. For LLMs equipped with normalization layers such as LayerNorm or RMSNorm, we find that changes in the direction of the target hidden state are essential for injecting new knowledge, whereas the magnitude contributes little. However, we also find that the norm of the target hidden state continuously accumulates during gradient descent optimization. We further show that this unnecessary norm growth increases interference with non-target knowledge, thereby degrading editing locality. Building on this insight, we propose Spherical Editing (SE), a simple yet effective method that constrains the target hidden state to remain on the sphere with the same radius as the original state throughout optimization. By eliminating unnecessary radial growth while retaining essential directional changes, SE reduces interference with other knowledge without sacrificing editing success. Extensive experiments on three LLMs, two benchmarks, and 10,000 edits demonstrate that SE consistently improves editing locality while maintaining editing effectiveness. Code is available at https://anonymous.4open.science/r/se-review-artifact-5869/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.