A Conditional Joint Matching of Response Profiles for LLM Unlearning
Abstract
Honouring a deletion request means making a language model behave like one retrained without the deleted facts, but retraining per request is too costly. Reference-based methods instead move each forgotten fact's likelihood to the level of never-seen facts, reducing the fact to one number. Yet memorisation also decouples a fact's probes, the phrasings that test it: on the TOFU benchmark, the correlation between direct and paraphrased likelihoods falls from 0.91 on never-seen facts to 0.15 on memorised ones, a drop that no order-preserving per-probe target can undo. Our intuition is that a forgotten fact should resemble comparable never-seen facts on all its probes at once. We propose conditional joint matching: each fact becomes a response profile of log-likelihoods under six probes, and entropic optimal transport couples each forgotten profile with never-seen profiles of similar relation, length and difficulty to build that profile's target. Unlike per-probe targets, these whole-profile targets carry the never-seen correlation, and training stops by a rule that needs no retrained model. On TOFU at 1B parameters, the method has the lowest mean per-fact distance to the retrained model among ten target assignments and four published objectives, 0.785 nats against 0.906 for unconditional per-probe matching, but only that comparison survives correction for multiple testing. Shuffling reference probes across facts within relation and difficulty strata gives a statistically indistinguishable 0.797, so the gain cannot come from how probes pair beyond those strata. Closer per-fact agreement, however, brings more leakage under sampling; the method also fails TOFU's Forget Quality test and trails a single-number target on a second model family. These results establish that forgetting must restore how probes co-vary across facts, not just a likelihood level, and that never-seen facts supply this structure but not the retrained model's level, so a retrain-free estimate of that level is the missing piece.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.