Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets
Abstract
Large language models often know the truth and say otherwise: a model that answers correctly when asked neutrally will affirm a user's mistaken belief, or misstate a fact its system prompt wants hidden, once the context rewards it. Such deception is not a piece of knowledge but a behavior conditioned on context, and current remedies treat it accordingly only in part: honesty and preference training suppress it on the training distribution, monitors catch it after the fact, and machine unlearning, the natural tool for removing a behavior from the weights, is built to forget facts that a deceptive model still needs. We propose to unlearn when a model deceives rather than what it knows, with a contrastive forget unit built from the model's own realized deceptions: the same question under a deception-triggering and a neutral context, admitted only where belief holds and behavior flips. Extending standard objectives to this unit exposes a dilemma. Suppression objectives such as NPO leave much of the deception in place. Target-based objectives, which distill the model's neutral behavior into the pressured context as context distillation and consistency training do, remove it but induce what we term context blindness: a target generated without the context teaches the model to stop reading it, eroding benign system-prompt instructions, secret-keeping and the reasoning a monitor would inspect, a failure that deception rates and capability benchmarks cannot see. We introduce PACT, which trains toward pressure-aware counterfactual targets (the model's own honest response, with a trace that registers the pressure and declines to be moved by it) while retaining the benign uses of the triggering context. On two 32B reasoning models, PACT reduces held-out deception from over to under while system-prompt adherence, secret-keeping and the reasoning trace stay at the base model's level. Scored on both sides of the trade-off, suppression retains but forgets little and distillation forgets but retains a third, while PACT is the only objective high on both, with a tug-of-war score of and against at most and for any baseline. Like removed knowledge, removed deception is shallow under relearning, and terms that simulate the attacker hold it only at a cost in context use.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.