AlignEdit: Alignment-Preserving Knowledge Editing for Post-Trained Language Models
Abstract
Knowledge editing updates specific knowledge while preserving unrelated behavior. This is challenging for post-trained language models, where edits must work with model-specific chat templates and open-ended generation without disrupting acquired capabilities. Yet most methods and evaluations target language models used directly after pretraining and assess fixed target continuations. We introduce an evaluation protocol combining model-specific input formatting with response-level correctness. It shows that representative methods struggle to balance Reliability, Generality, and Locality during sequential editing, even with chat-formatted editing inputs. We then propose AlignEdit, which independently constructs knowledge-specific teachers from the original model through localized fine-tuning and distills them into a shared student using dense supervision along student-generated trajectories. We evaluate AlignEdit on Qwen3-4B and SmolLM3-3B using CounterFact and ZsRE, together with five general-capability benchmarks. Across both editing benchmarks, AlignEdit achieves a better balance among Reliability, Generality, and Locality than the evaluated baselines, with limited degradation in average general-capability performance. These results underscore the importance of accounting for post-training behavior in practical knowledge editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.