Crossing Lethe: Steering Probability Flows for Data Unlearning
Abstract
Data unlearning (the targeted removal of specific training examples from a deployed model) is increasingly important for privacy compliance, user rights, and safe model updating. While unlearning has been actively studied for energy-based models and diffusion models, comparatively little work addresses flow matching models, in part because it is unclear how removing data should alter the learned dynamics. Flow matching trains continuous generative models by learning a time-dependent velocity field whose estimates aggregate contributions from all training samples, making post-training selective removal nontrivial. We propose a flow decomposition framework that adapts to the level of data access and operational constraints. We first introduce Loss-level Flow Decomposition (LFD), that explicitly subtracting the forget-set contribution. Importantly, LFD serves as a global optimizer that obviates the need to retrain the model. LFD provides high fidelity in standard settings where both retain and forget datasets are available. We then investigate privacy-restricted scenarios where explicit access to forget samples is prohibited or governed by strict constraints. Specifically, we introduce an Energy-Based Reformulation (EBR) that leverages a proxy classifier rather than direct forget data samples. We rigorously analyze both these approaches. We also validate our framework on several datasets, demonstrating that these methods effectively balance unlearning performance with the preservation of model fidelity across both standard and restricted data regimes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.