acceptodds
Under review as a conference paper at ICLR 2027

SAVER: Sequential Anytime-Valid Editing Risk-control in Large Language Models

Abstract

Sequential knowledge editing (KE) in large language models (LLMs) can accumulate interference, weakening both edits success and preservation of reference behavior. Existing editing methods improve update quality and reduce interference, but failures in individual updates can still go undetected. Once committed, these updates alter the subsequent edits, allowing their effects to persist and accumulate. To address this, we propose SAVER (Sequential Anytime-Valid Editing Risk-control), a model- and editor-agnostic pre-commit controller for sequential KE. Specifically, SAVER measures prediction-set miscoverage on generality and locality probes and uses sequential risk evidence to select the smallest admissible boundary and guide edit admission. We further develop interference-aware sampling with a control-variate risk estimator to reduce full-probe evaluation cost. Moreover, we establish anytime-valid false-alarm control across time and candidate boundaries, together with detection-delay bounds under persistent excess risk and the stated monitoring conditions. Extensive experiments across editors, backbone LLMs, and datasets demonstrate improved reference preservation while retaining useful updates and reducing monitoring cost. Our code is available at https://anonymous.4open.science/r/SAVER-E7D6/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.