ReSAT: Training LLMs Resistant against Refusal Ablation Attacks
Abstract
Open-weight language models allow downstream users to modify internal com- putations as well as prompts, creating a distinct challenge for preserving learned safeguards. We investigate whether safety training protects refusal and safety against the removal of multiple activation directions, rather than just one. We introduce Refusal Subspace Ablation, an attack that selects complementary di- rections and jointly removes their span throughout generation without updating model weights. Starting from one direction, it adds components outside the already selected span, choosing them by their effect under joint ablation. Our evaluations show that some current existing safety-trained models resist a selected single- direction attack yet remain vulnerable to joint ablation. To narrow this gap, we propose Refusal Subspace Ablation Training: mixed supervised fine-tuning fol- lowed by safety-training rounds that supervise safe responses under joint ablation and refresh candidate directions between rounds. We evaluate ReSAT on Llama- 3-8B, Mistral-7B, Qwen2.5-14B, against RSA and other refusal-ablation attacks, with jailbreak attacks and utility test. In key settings, ReSAT achieves lower attack success rates than most other baselines while preserving clean MMLU accuracy. By connecting adaptive attack evaluation with safety training, our work offers a practical path toward more robust safeguards for open-weight LLMs
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.