Safety Alignment of Large Language Models via Intrinsic Sparse States
Abstract
Safety alignment requires models to reject harmful requests without refusing legitimate requests on sensitive topics. Response-level supervision specifies desired answers but leaves the internal distinction between these requests to be learned indirectly. We introduce ISA, an **I**ntrinsic **S**parse-state **A**lignment framework that complements response-level training with explicit supervision of sparse autoencoder (SAE) feature activations at the final prompt token. A trained probe identifies predictive features, and intervention tests retain those that improve harmful-request safety within a benign-refusal budget. The selected features then support margin losses on prompt-risk scores and bounded activation guidance during LoRA adaptation. Across Llama-3.1-8B, Qwen3-8B, and Gemma-2-9B, ISA reduces unsafe-response rates on three harmful-request benchmarks and refusal rates on two benign-sensitive benchmarks relative to SFT trained on the same response targets. ISA substantially reduces benign refusal compared with SFT+DPO, although DPO achieves lower unsafe-response rates on two benchmarks. Construction controls further show that behavioral screening reduces both harmful compliance and benign refusal on XSTest relative to predictive selection alone. Loss ablations distinguish how the internal objectives contribute to this safety–refusal balance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.