Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning
Abstract
Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of , yet storage identity agrees with the better intervention on only targets, while low-rank adaptation (LoRA) wins . We introduce Intervention Score, which ranks editable groups by the predicted effect of the actual unlearning update while accounting for collateral damage, and use it to form the static intervention-value baseline (). We then introduce selective dynamic intervention re-ranking (), which revisits that subset only when a calibrated probe justifies the comparison. On the Natural-TOFU dataset, our method has positive descriptive margins in comparisons between methods and objectives, although several are near zero. On the LACUNA localization-precision benchmark, our mean terminal utility is higher in all six negative preference optimization (NPO) and SimNPO comparisons: NPO margins range from to , and SimNPO margins range from to . The gradient-difference (GradDiff) objective reveals substantial field dependence. Relative to , the primary four-field GradDiff evaluation has six wins, six ties, and no losses, with mean and median paired gains of and . The evidence supports separating localization, initial intervention selection, and checkpoint-dependent support revision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.