acceptodds
Under review as a conference paper at ICLR 2027

A Window Is Not a Door: Mapping the Limits and Conditions of Knowledge Editing in Language Models

Abstract

Interpretability makes factual knowledge in language models observable, but observability does not imply editability. We identify a gap between locating factual signals and reliably changing knowledge, defined by control, transfer, and collateral damage. In the tested settings, localization-derived interventions move target answers without achieving selective control within the specified collateral budgets. Post-edit target hits likewise conceal already-correct answers, limited transfer, and damage to other knowledge. Reading out information and eliciting a target answer are therefore insufficient evidence of reliable knowledge editing. This gap is not uniform: the model’s pre-edit support for the target answer helps organize where correction succeeds. Targets near the front of the base distribution, but not initially generated correctly, form a region of correction opportunities. Candidate rank predicts correction yield in independent fixed-budget tests, while far-rank successes reveal a model-dependent boundary rather than a hard cutoff. Context and weight interventions further expose distinct routes to changing answers, with different benefits and costs. Together, these findings shift the central question of knowledge editing from where knowledge resides to which facts can be changed, through which intervention, and at what cost. Interpretability provides a window into knowledge; a door to reliable editing depends on existing candidate support and the conditions of intervention.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.