acceptodds
Under review as a conference paper at ICLR 2027

NormGuard: Evidence-to-Norm Reasoning for Dynamic Personalized Safety

Abstract

Personalized safety requires AI assistants to adapt their behavior to user-specific risks and constraints. Existing work, however, largely assumes that safety-relevant user context is available upfront or remains static. In real-world dialogues, such information is often revealed progressively across turns, requiring assistants to continuously accumulate evidence about the user. Beyond tracking this evolving context, a second challenge is determining how the observed evidence should affect the assistant’s behavior, as existing context- or profile-conditioned approaches typically leave this mapping implicit. To study these challenges under controlled information asymmetry, we introduce a dynamic personalized-safety benchmark in which safety-relevant user information is progressively revealed while the complete user profile remains hidden from the assistant. We further propose NormGuard, an evidence-to-norm framework that separates user-state estimation from behavioral constraint formation. NormGuard progressively constructs a perceived profile from observable dialogue evidence and induces structured, evidence-grounded safety norms that specify applicable constraints, priorities, and supporting evidence for response generation. Extensive experiments show that NormGuard improves personalized safety over competing baselines, while ablation studies demonstrate complementary benefits from perceived-profile tracking and explicit safety-norm guidance. These results show that personalized safety benefits from explicitly translating progressively revealed user evidence into behavioral constraints, rather than relying on profile information alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.