Beyond Aggregate Forgetting: Knowledge Graph Unlearning Can Disagree with Retraining
Abstract
Knowledge-graph unlearning (KGU) is typically evaluated by how closely an updated model matches retained-data retraining on aggregate ranking metrics such as mean reciprocal rank (MRR). We show that this criterion can accept models whose individual predictions disagree with retraining far more than two independent retrains of the same retained graph, REFA and REFB, disagree with each other. To isolate this failure mode, we study a restart-and-repair (R+L) setting in which deletion causes complete support loss, removing all positive training evidence for selected entities while their embeddings remain in the model. R+L reinitializes these unsupported entities before retained-data repair. This brings forgotten-set MRR close to REFA while preserving test performance, yet the agreement does not extend to individual predictions. Across 36 deletion settings on three graphs, each removing roughly 10% of the training facts, R+L leaves an MRR gap of.008 against a.004 retrain-to-retrain null, yet exceeds that null in log-rank in every one, at a mean calibrated excess of.256. The discrepancy is not specific to R+L: on matched random-fact deletion requests, Finetune, GraphDPO and SGU all disagree with REFA more than REFB does, in every setting. We then ask what drives it. Across 72 configurations spanning six scoring families, the strongest effect appears in the bilinear ones, and only when an affected entity is ranked as a candidate rather than supplied as query context; a matched background analysis shows that restarted entities are systematically promoted as candidates even in unrelated contexts. An oracle diagnostic that rescales only their norms to the reference removes 92.2% of the excess, identifying embedding-scale mismatch as the dominant cause, but it consumes a retraining reference. We therefore derive from it Retained-Norm Rescaling (RNR), a correction that uses only the updated model and the retained graph; REFA and REFB are used only for evaluation. RNR removes 75.6% of the calibrated log-rank excess, at no cost in test performance. Under a stronger training recipe, the excess survives on two of the three graphs. Throughout, we judge an unlearning update not by how closely its aggregate metrics match retraining, but by whether its individual predictions disagree with a retrained model more than two retrained models disagree with each other.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.