Inference-Time Machine Unlearning via Gated Activation Redirection
Abstract
Large Language Models (LLMs) memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety. Machine unlearning seeks to remove the influence of a targeted forget set D_f while preserving model performance, ideally approximating a model retrained from scratch without D_f. Once an LLM is in use, every new request to make it forget specific content demands updating its weights. However, unlearning through parameter updates is expensive, hard to audit, and can be undone by quantization. We show that unlearning can be enforced entirely at inference time, without training, gradients, or weight changes. We introduce Inference-Time Unlearning via Gated Activation Redirection (GUARD-IT), a training- and gradient-free method that unlearns via input-dependent activation steering at inference time. guard stores the content to be forgotten as a small library of activation directions, and during inference, it routes each query through a similarity gate that activates only for relevant directions and applies them as a norm-preserving rotation of the residual stream, while unrelated queries pass through the unmodified model. The same design carries across three model families and nine checkpoints from 0.8B to 8B parameters, and new forget requests are absorbed by one offline pass of forward passes. On TOFU, against 16 gradient-based and inference-time baselines, and on MUSE and WMDP, GUARD-IT forgets without breaking the model, and on TOFU it is the only method that suppresses memorization in every Llama configuration without collapsing: it keeps utility and fluent generation in every configuration, moves the model's output distribution closest to a model that never saw the forgotten data on the forget01 split, survives 4- and 8-bit quantization and ten sequential forget requests, and holds under jailbreak attacks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.