acceptodds
Under review as a conference paper at ICLR 2027

CINDER: Twin-Referenced Training for Relearning-Resistant Unlearning

Abstract

LLM unlearning methods that pass their own evaluations are routinely undone by a few steps of fine-tuning on part of the forget set. We argue the right reference for resistance is not "unrecoverable" but a retain-only twin — a model trained without the forget data — and train against it directly. CINDER simulates the attacker's own attack class inside training (full-parameter fine-tuning on a random half of the forget facts, first-order meta-gradient) and penalizes recovery of the held-out half through a softplus penalty anchored at the twin's post-attack floor. We evaluate with the hypotheses the right way round: method-level non-inferiority to the twin at a pre-specified margin — the spread between two independent retrains — across corpora, with every axis required. On sixteen corpus seeds at 0.5B, CINDER's post-attack recovery exceeds the twin's by 0.08 nats (one-sided 95% bound 0.21, margin 0.36), 1.17 nats below NPO (95% CI 0.94–1.39); on sixteen fresh corpora generated after the analysis was fixed, +0.16 nats (bound 0.29), 1.16 below NPO. It is non-inferior on six of eight axes; on the circuit criterion and utility non-inferiority is not shown. A frozen, preregistered eight-axis protocol, used as a screen, agrees: relearning is not rejected on 15 of 16 seeds (16 of 16 fresh) against at most 4 for five baseline families, and no axis on 8 of 16 (3 of 16 fresh, C1 rejecting eleven) against at most 2. Ablations over eight to sixteen seeds show that the floor, the full-parameter inner attack and the fact-level split are each needed, that a floor taken from the pretrained base — no twin at training time — works as well as the twin's, and that adding our floor to MUDMAN brings it within the margin. At 1.5B and 3B relearning is not rejected on any of twelve seeds, but utility is paid for. We report what this does not show: a pass of a finite battery bounds nothing, the objective overshoots its floor, and a twin, though not needed for training at 0.5B, is still needed to evaluate.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.