acceptodds
Under review as a conference paper at ICLR 2027

Fine-tuning Disrupts Edit Readout: A Mechanistic Perspective on Model Editing Robustness

Abstract

Model editing provides a precise way to update the parametric knowledge of Large Language Models (LLMs). However, it has been shown that fine-tuning of edited models for downstream adaptation will erase the intended edits in existing works. This paper studies this problem through the lens of layer-wise information flow in Transformers. Our systematic layer-wise ablation study reveals that edited memory is mainly anchored in shallow editing layers, while fine-tuning primarily disrupts the deep editing layers and adjacent post-edit layers, jointly forming the readout circuitry that decodes this memory into the final prediction. As a result, reverting this readout circuitry alone is often sufficient to recover editing success. We further verify this layer-wise information-flow perspective through null-space analysis and attention-based visualization. Motivated by this key finding, we propose Layer-wise Model Interpolation (LMI), which selectively restores the disrupted readout circuitry to preserve edited knowledge while retaining downstream adaptation. LMI interpolates between the edited and fine-tuned model weights with layer-aware coefficients, up-weighting the edited model in the readout circuitry and the fine-tuned model elsewhere. Across four public datasets, three language models, and two editing methods, LMI achieves the best balance between edit preservation and downstream utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.