Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution
Abstract
Standard unlearning evaluations measure behavioral suppression in full precision, immediately after training, despite quantization being a common deployment transformation for language models. Prior works have shown that 4-bit post-training quantization can reverse machine unlearning. We show this is a systematic dual failure. Gradient-based methods that achieve meaningful forgetting lose it under compression, while methods that survive quantization barely change the model. Both failures have the same root cause. Across all baselines, per-parameter RMS updates are 47 to 828× below the canonical NF4-calibrated floor used in our implementation, so updates diffused across billions of parameters cannot clear quantization bin boundaries, a consequence we formalize as a **sparsity-permanence tradeoff**. We present **MANSU** (**M**echanistic-**A**ligned **N**ull-**S**pace **U**nlearning), which addresses this tradeoff by treating it as an update-budget allocation problem. Causal circuit attribution isolates a compact, causally implicated forget circuit, a diagonal-Fisher mask restricts updates to retain-safe coordinates within that circuit, and a quantizer-calibrated magnitude floor, applied post-hoc, makes the resulting update deployment-visible, which we verify directly on the trained checkpoints. Matched controls (the same floor applied to a random circuit, an inverse circuit, a globally projected update, and a Fisher-ratio selective baseline) show that neither circuit concentration nor magnitude flooring alone reproduces this behavior in the reported ablations. We additionally introduce **Circuit Attribution Divergence (CAD)**, a mechanistic verification metric providing evidence of circuit-level structural change beyond behavioral suppression, corroborated by independent linear-probe and activation-patching analyses. Across multiple model families and WMDP hazard domains, MANSU consistently yields a non-positive NF4 PTQ gap while retaining substantially higher general capability than aggressively optimized baselines, whereas gradient ascent, for instance, partially reverses under compression (BF16 forget 0.260 to NF4 0.310, PTQ gap +0.050). We scope our boundary-crossing analysis to NF4 and report GPTQ and INT8 results separately as empirical transfer checks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.