acceptodds
Under review as a conference paper at ICLR 2027

A Few Bits Break Forgetting: Bit-Flip Relearning Attack against Machine Unlearning

Abstract

Machine unlearning (MU) is an emerging paradigm to remove specified knowledge from trained models. Despite significant progress, recent studies reveal that unlearned knowledge can be recovered through relearning attacks, raising concerns about MU’s robustness. However, existing relearning attacks typically rely on fine-tuning and model update capability for dense parameter modifications, making them detectable by deployment-time integrity verification and less practical. This raises a critical question: can relearning attacks be achieved through only minimal runtime bit-level modifications? As bit-level updates are sparse and discrete, two challenges arise: balancing relearning with utility and efficiently identifying effective bit flips. Existing bit-flip attacks mainly degrade model behavior, which fundamentally differs from relearning and fails to address the challenges. In this paper, we propose **BFRA**, a **B**it **F**lip **R**elearning **A**ttack that is applicable across tasks and MU methods, enabling fine-tuning-free relearning through minimal bit flips. To balance relearning and utility, BFRA formulates relearning as a constrained discrete optimization problem that maximizes forgotten knowledge recovery while constraining retain utility drop. To improve efficiency, BFRA adopts a progressive bit-flip strategy combining First-Order Candidate Selection and Trial-Loss Bit Picking. Extensive experiments across 3 tasks, 14 model-dataset settings, and 12 MU methods show that BFRA substantially recover forgotten knowledge with only a few dozen bit flips while keeping high retain utility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.