acceptodds
Under review as a conference paper at ICLR 2027

RExMap: Counterfactual Recovery for Model-Edit Rollback after Composition and Serving Drift

Abstract

A revoked edit can disappear from the live weights while its influence survives inside every later update optimized in its presence. RExMap recovers the corresponding counterfactual history: its ledger unwinds the retained suffix, removes the revoked update, re-estimates each retained request on the corrected state, and evaluates the recovered model on the edited fact, 500 clean prompts, 200 attacked probes, and an off-target entity neighborhood. In 24,000 Llama-3.1-8B transactions, exact all-conditions restoration falls from 97.4% at depth 1 to 82.8% at depth 20 in FP32, 68.7% under INT8, and 60.9% with dual overlap and non-LIFO revocation. Re-estimating retained descendants recovers 12.0 points over direct subtraction at depth 20, identifying state-dependent descendants as a correctable source of rollback error. A transaction-state score predicts behavioral validity with 0.914 AUROC and a 2.7% false-valid rate; routing uncertain cases through checkpoint replay raises restoration by 22.0 points over uniform subtraction to 97.8%, with a 60.1% direct fast path and 0.31 s median rollback. The depth and serving-representation ordering recurs across a reversible ROME-style writer, Mistral-7B-v0.3, and MQuAKE. RExMap therefore turns reversibility from a write-time algebraic property into a behaviorally testable recovery decision for evolving edited models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.