Turning Abliteration into Model Self-Destruction in Open-Weight LLMs
Abstract
Open-weight LLMs can be easily un-censored through abliteration, which is a method that works by removing the refusal direction from the residual stream. In this paper, we contribute an abliteration defense that makes the attack backfire. We ensure removing the refusal direction leads to either model collapse or persistent refusal, while leaving the clean model intact. During alignment we supervise the model on simulated abliteration in random subspaces of the refusal direction, training the model to emit a collapse token with strong abliteration and maintain refusal under gentle partial abliteration. On Llama-3.1-8B, abliteration attack-success (HarmBench) falls from 0.72 to 0.00, and we show the effect generalizes across multiple model families and to out-of-distribution and capability-preserving adaptive abliteration. We also provide a discussion of the field, as well as future work that can build off of our method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.