Forget What Comes Next: LLM Unlearning via Statistical Independence
Abstract
Large language models (LLMs) can memorize and reproduce fragments of their training data, raising privacy and legal concerns that motivate machine unlearning. Yet poisoned training examples can remain statistically detectable after unlearning, highlighting the need for a deeper theoretical understanding of how unlearning objectives relate to memorization. In this work, we introduce PINT, a novel algorithm for machine unlearning that uses the Hilbert-Schmidt Independence Criterion (HSIC) to reduce the statistical dependence between tokens to be forgotten and the internal representations used to predict them. We formally relate our HSIC-based objective to exact memorization through a computable upper bound, connecting the training objective directly to the model's ability to reproduce forgotten tokens. In extensive experiments, we compare forgetting-utility trade-offs across nine settings with models up to 7B parameters and demonstrate scalability to 70B parameter models. Our evaluations cover unlearning personal information, hazardous knowledge in biology and cybersecurity, copyrighted books and news articles, and Gaussian poisons. PINT achieves the strongest forgetting-utility trade-offs on most benchmarks. On Gaussian poisons, it is the only method to make the poison undetectable while maintaining near-original model utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.