Localized Recall-Capped Unlearning for Diffusion Language Models
Abstract
Diffusion language models (DLMs) are emerging as an alternative to autoregressive language models. Unlearning is needed to remove unwanted knowledge from trained models, yet weight-space unlearning in DLMs remains largely unexplored. Empirically, existing objectives either suppress target knowledge at the expense of general capability or preserve utility with insufficient forgetting. We identify update location and gradient control as two key factors underlying these failures. Our method, Localized Recall-Capped Unlearning (LRCU), uses causal-mediation localization to select three consecutive transformer blocks. It updates only these blocks using capped per-token cross-entropy, with each token contributing zero forgetting gradient once its cross-entropy exceeds the cap. LRCU requires no reference model and updates less than 10% of parameters, with a 1.5 average per-step speedup over the compared full-model baselines on LLaDA. Across WMDP-bio, WMDP-cyber, RWKU, and TOFU on two DLM backbones, LRCU achieves a competitive forgetting-utility trade-off compared with existing methods. These results highlight the importance of update location and gradient control for effective DLM unlearning. We will release an Open-DLU codebase upon acceptance for weight-space unlearning in DLMs, together with unlearned checkpoints to support further research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.