Fine-tuning Disrupts Edit Readout: A Mechanistic Perspective on Model Editing Robustness
Abstract
Model editing provides a precise way to update the parametric knowledge of Large Language Models (LLMs). However, it has been shown that fine-tuning of edited models for downstream adaptation will erase the intended edits in existing works. This paper studies this problem through the lens of layer-wise information flow in Transformers. Our systematic layer-wise ablation study reveals that edited memory is mainly anchored in shallow editing layers, while fine-tuning primarily disrupts the deep editing layers and adjacent post-edit layers, jointly forming the readout circuitry that decodes this memory into the final prediction. As a result, reverting this readout circuitry alone is often sufficient to recover editing success. We further verify this layer-wise information-flow perspective through null-space analysis and attention-based visualization. Motivated by this key finding, we propose Layer-wise Model Interpolation (LMI), which selectively restores the disrupted readout circuitry to preserve edited knowledge while retaining downstream adaptation. LMI interpolates between the edited and fine-tuned model weights with layer-aware coefficients, up-weighting the edited model in the readout circuitry and the fine-tuned model elsewhere. Across four public datasets, three language models, and two editing methods, LMI achieves the best balance between edit preservation and downstream utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.