SafetyFlip: Learning LLM Safety Boundaries from Bidirectional Counterfactual Instructions
Abstract
Balancing helpfulness and safety in large language models (LLMs) remains challenging: imprecise safe/unsafe boundaries can cause both unsafe compliance and over-refusal of safe but sensitive requests. We propose **SafetyFlip**, a data-centric framework that learns LLM safety boundaries from bidirectional counterfactual instructions. SafetyFlip constructs semantically matched SafeUnsafe pairs with structured boundary annotations identifying the preserved non-safety semantic frame, the safety-critical factor, and the required policy behavior. These pairs are used for **Boundary-Constrained Fine-Tuning (BCFT)**, which combines supervised fine-tuning with Safety Contrastive Regularization (SCR) to encourage safety-relevant representation shifts while preserving shared non-safety utility. On Qwen2.5-7B, SafetyFlip reduces HarmBench attack success from 58.4% to **6.5%** while keeping XSTest over-refusal low at **2.1%** and preserving MT-Bench utility. Our work offers a scalable, low-supervision pathway for boundary-aware alignment without inference-time overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.