acceptodds
Under review as a conference paper at ICLR 2027

Empathetic yet Guarded: Dynamic Personalized Safety Alignment in Multi-Turn Conversations

Abstract

While current safety alignment methods have significantly improved the general harmlessness of Large Language Models (LLMs), they predominantly enforce uniform safety boundaries and overlook the complex, individualized contexts of human interactions. This uniform approach often leads to critical failures. For example, models may provide advice that is universally benign but harmful to specific vulnerable users. Alternatively, they may excessively cater to user personas, risking a bypass of general safety rules. To bridge this critical gap, we introduce MDPSA, a novel Multi-Turn Dialogue Dataset for Personalized Safety Alignment. Our dataset spans seven high-risk domains and comprises 58 fine-grained response strategies, uniquely capturing the dynamic evolution of user emotional states and complex, context-driven jailbreak scenarios. To validate its effectiveness, we conduct comprehensive experiments across various post-training and alignment paradigms. Empirical results demonstrate that while standard models struggle heavily with personalized safety, training on MDPSA effectively empowers them to resolve the tension between general safety and individual adaptability, achieving significant performance gains. Our work formally establishes a robust data foundation and empirical baseline for building safe, empathetic, and context-aware LLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.