acceptodds
Under review as a conference paper at ICLR 2027

Sensitivity Awareness in Multi-Turn Interactions: An Information-Theoretic Perspective

Abstract

Despite their impressive capabilities, large language models (LLMs) remain vulnerable to many adversarial attacks. A key weakness is their lack of sensitivity awareness (SenA), i.e., their inability to consistently enforce access rights when processing and sharing sensitive data in multi-turn interactions. Although this problem has recently gained attention, no unified framework exists for analyzing SenA over multiple turns. We address this by defining hop information leakage, a per-turn measure of information disclosure based on mutual information that composes exactly across turns via the chain rule. Using this notion, we derive an upper bound on an adversary's success probability, directly connecting cumulative information leakage to the attacker's advantage through average conditional min-entropy; empirically demonstrate how an uninformed LLM system can leak sensitive information even if it never states the target value outright; propose BayWatch, an information-theoretic SenA guardrail that reduces the attacker's success rate by up to 45 percentage points compared to the state-of-the-art defense. These contributions offer a principled, empirically vetted basis for understanding the capabilities and fundamental limits of multi-turn adversaries against sensitivity-aware systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.