acceptodds
Under review as a conference paper at ICLR 2027

SafetyFlip: Learning LLM Safety Boundaries from Bidirectional Counterfactual Instructions

Abstract

Balancing helpfulness and safety in large language models (LLMs) remains challenging: imprecise safe/unsafe boundaries can cause both unsafe compliance and over-refusal of safe but sensitive requests. We propose **SafetyFlip**, a data-centric framework that learns LLM safety boundaries from bidirectional counterfactual instructions. SafetyFlip constructs semantically matched SafeUnsafe pairs with structured boundary annotations identifying the preserved non-safety semantic frame, the safety-critical factor, and the required policy behavior. These pairs are used for **Boundary-Constrained Fine-Tuning (BCFT)**, which combines supervised fine-tuning with Safety Contrastive Regularization (SCR) to encourage safety-relevant representation shifts while preserving shared non-safety utility. On Qwen2.5-7B, SafetyFlip reduces HarmBench attack success from 58.4% to **6.5%** while keeping XSTest over-refusal low at **2.1%** and preserving MT-Bench utility. Our work offers a scalable, low-supervision pathway for boundary-aware alignment without inference-time overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.