EAGPO: Edit-Anchored Generalization Policy Optimization for Knowledge Editing
Abstract
Knowledge editing aims to precisely update factual knowledge in large language models. However, existing methods struggle to generalize the editing knowledge to semantic variants, limiting their practical application. This fragility often arises because these methods optimize solely for a single target expression, neglecting the broader semantic neighborhood that real-world queries inhabit. To bridge this gap, we propose EAGPO, a novel framework that enhances knowledge editing generalization to semantic variants through Edit-Anchored Generalization Policy Optimization. It first employs a standard knowledge editing method to precisely integrate the new knowledge. Then, it performs reinforcement learning training via two core components: an Anchor-Generalization Training (AGT) mechanism that optimizes over structured query clusters, and a Policy Gradient Alignment Advantage (PGAA) module that stabilizes updates through geometric gradient alignment. Theoretically, EAGPO strengthens knowledge protection, provides a tighter generalization error bound, and reduces policy gradient variance. Empirical results on the KnowEdit benchmark show that EAGPO significantly outperforms strong baselines across three types generalization tasks, improving generalization performance by 25%–35% over the original knowledge editing methods, while achieving a 25% improvement in computational efficiency over GRPO training. Our code is available at https://anonymous.4open.science/r/Edit-GRPO-119DEAGPO https://anonymous.4open.science/r/Edit-GRPO-119D.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.