acceptodds
Under review as a conference paper at ICLR 2027

Gradient as Behavior Fingerprint: Stable Sequential Model Editing via Gradient Drift Regularization

Abstract

Large language models (LLMs) have achieved remarkable success across a wide range of tasks, but outdated or incorrect knowledge may persist in their parameters, leading to hallucinations and factual errors. Model editing aims to address this issue by enabling targeted updates to specific knowledge. However, in sequential editing settings, repeated updates can interfere with one another, resulting in accumulated instability and degradation of previously edited knowledge. In this paper, we revisit sequential editing from a gradient-based perspective. We use the gradient of the target log-likelihood with respect to the editing residual as a target-conditioned descriptor of first-order local perturbation response. We establish a Taylor bound showing that the gradient discrepancy between two state-specific reference centers controls the difference between their centered local responses up to a second-order remainder. Building on this diagnostic, we propose Gradient Drift Regularization (GDR). For each historical edit, GDR stores a gradient anchor at its insertion-time reference center and later computes a query gradient from the current model state when that edit is selected. GDR uses the discrepancy between these gradients to construct a fixed, dimension-wise quadratic penalty on the current editing residual, discouraging updates along coordinates with larger observed historical discrepancies. To improve efficiency, we further introduce a preconditioned key-space selection mechanism that identifies historical edits potentially affected by the current update, limiting query-gradient evaluation to a top- subset. Experiments on LLaMA-3.1, GPT-2-XL, and Qwen-2.5 show that GDR preserves overall model performance and improves the retention of historical edits over long editing sequences. The results also reveal a strong empirical association between lower gradient drift and better historical retention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.