Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
Abstract
Pretrained large language models (LLMs) achieve remarkable performance across natural language tasks, yet their internal knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated knowledge. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject the passage but fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that because the injected knowledge is novel, the pre-edited model's own rollouts rarely cover it, limiting the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios. We release our datasets and evaluation code at https://anonymous.4open.science/r/hpse-review-C778.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.