acceptodds
Under review as a conference paper at ICLR 2027

Striking the Harmonization: Personalized Safety via the Consistency Constraints

Abstract

Recent research on personalized safety in large language models has separately investigated two phenomena—personalized jailbreaking and personalized maladaptation—yet these lines of work remain disjointed, offering no unified account of the fundamental tension between general safety constraints and individual-specific requirements. We propose a Consistency-based framework that reconceptualizes both safety and personalization as two forms of contextual consistency: the former internalized as an intrinsic model property, the latter grounded in input context. Under this unified consistency perspective, we formalize personalized safety as a Constrained Reinforcement Learning problem. Specifically, we minimize safety consistency to strictly enforce general safety boundaries, while maximizing the personalized reward (the negation of personalized consistency) to drive individual preference satisfaction. This formulation recasts the apparent tension not as a trade-off, but as a jointly satisfiable condition under explicit consistency constraints on safety-personalization feasibility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.