Exact, Reversible, and Verifiable Lifecycle Control for Routed Parametric Memory
Abstract
Deleting, correcting, or reverting what a deployed language model knows is today an optimisation problem on entangled weights. Unlearning objectives and model editors move parameters until a behavioural probe is satisfied. The result is approximate, depends on the order of operations, and can be verified by no one but the operator. The problem disappears when knowledge is placed rather than optimised. Block-Rollback stores injected knowledge in the content-routed memory blocks of a hierarchical-memory language model whose backbone, the anchor, stays frozen. Every block is a deterministic function of the documents it holds and a public seed. Insert, edit, delete, and revert therefore reduce to rebuilding one block from an append-only log, and each operation lands element-wise on the counterfactual model: the one that never saw the deleted fact, only ever saw the corrected fact, or stopped at the reverted step. We test this on five public anchors from GPT-2-XL to Llama-2-7B and on two anchors pretrained from scratch with the memory pathway, across the KnowEdit, TOFU, and MUSE benchmarks. More than 17000 locality checks find no violation, and every rebuilt block matches its counterfactual exactly. On the same base, gradient editors halve their locality within ten edits and cannot revert, and optimisation-based unlearners remain separable from the counterfactual by membership, extraction, and jailbreak probes. Block-Rollback matches it on every probe, including after int8 and NF4 quantisation. A third-party certificate rejects 10000 forged deletions at the cost of rebuilding one block. Per-operation cost tracks the touched block rather than the corpus or the record's age. A read-only routing gate keeps general capability at the anchor's values through 300 edits.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.