acceptodds
Under review as a conference paper at ICLR 2027

AlignEdit: Alignment-Preserving Knowledge Editing for Post-Trained Language Models

Abstract

Knowledge editing updates specific knowledge while preserving unrelated behavior. This is challenging for post-trained language models, where edits must work with model-specific chat templates and open-ended generation without disrupting acquired capabilities. Yet most methods and evaluations target language models used directly after pretraining and assess fixed target continuations. We introduce an evaluation protocol combining model-specific input formatting with response-level correctness. It shows that representative methods struggle to balance Reliability, Generality, and Locality during sequential editing, even with chat-formatted editing inputs. We then propose AlignEdit, which independently constructs knowledge-specific teachers from the original model through localized fine-tuning and distills them into a shared student using dense supervision along student-generated trajectories. We evaluate AlignEdit on Qwen3-4B and SmolLM3-3B using CounterFact and ZsRE, together with five general-capability benchmarks. Across both editing benchmarks, AlignEdit achieves a better balance among Reliability, Generality, and Locality than the evaluated baselines, with limited degradation in average general-capability performance. These results underscore the importance of accounting for post-training behavior in practical knowledge editing.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.