acceptodds
Under review as a conference paper at ICLR 2027

From Fact Overwriting to Causal Editing: Knowledge Evolution via On-Policy Self-Distillation

Abstract

While Knowledge Editing (KE) enables efficient updates, its dominant *Static Fact Overwriting* paradigm treats LLMs as discrete databases, forcibly injecting isolated facts. Clashing with un-evolved legacy priors, this triggers **Epistemic Dissonance**—a pathology where obsolete beliefs compel the model to explicitly negate the injected update. Idealized interventions reveal that this is a systemic paradigm limitation rather than mere algorithmic noise, with a zero-distortion stress test yielding a catastrophic 95.6% self-refutation rate. Given the causally driven nature of real-world knowledge, grounding updates in explicit causal narratives effectively collapses this conflict rate to just 6.6%, underscoring the imperative for **a paradigm shift toward Causal Editing**. To this end, we propose **CODE** (**C**ausal **O**n-policy **D**istillation for **E**diting), a two-stage causal editing framework that effectively integrates such causal logic into model parameters by coupling causal bootstrapping with asymmetric on-policy distillation. Experiments on LLaMA-3.1 and Qwen-2.5 show CODE suppresses self-refutation to 1.8% while securing robust multi-hop accuracy (up to 83.5%), with further probing demonstrating genuine causal internalization over surface memorization—elevating discrete fact overwriting into coherent, globally consistent knowledge evolution. Code is available [here](https://anonymous.4open.science/r/CODE-CD3E/).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.