Crush the Exposed Dark Side: Unlearning via Suppressing LLMs' Hazardous Output Tendency
Abstract
Large Language Models (LLMs) are widely deployed in real-world applications, yet their safety remains vulnerable to evolving jailbreak attacks. Existing defenses, including alignment and unlearning, typically rely on training data organized as question-answer pairs. As a result, their robustness is inherently bounded by the coverage of harmful prompts observed during training. Since natural language admits virtually unlimited reformulations, exhaustively covering the question space is infeasible in practice. To address this limitation, we propose Answer-Centric Unlearning (ACU), a framework that suppresses harmful answers exposed by the model itself. Specifically, ACU exposes latent harmful trajectories, iteratively suppresses them without conditioning on questions, and then trains on a retain set to restore normal utility. Experiments across four 8B backbones and one 32B model show that ACU improves robustness on eight additional jailbreak attacks while maintaining acceptable utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.