LLMs Already Know What Not to Share: Span-Level Contextual Integrity via Probing and Attention Masking
Abstract
As large language model (LLM) agents act on users' behalf, they are increasingly given access to personal information. Yet access does not imply permission to disclose: under Contextual Integrity (CI), the appropriateness of each information span depends on the current task. Given their large-scale training, we hypothesize that LLMs possess an implicit capacity for CI, reflected in contextual privacy signals encoded during prompt processing. To formalize this capacity, we propose the Contextual Linear Privacy Hypothesis, which models privacy-relevant information as a linear component of the span representation whose strength varies with context, and characterizes how changing context shifts the representation of a fixed span. We operationalize this hypothesis with layer-specific linear probes and introduce probe-guided attention masking, which prevents generated tokens from attending to cached positions classified as restricted without modifying the base model or generating additional privacy reasoning. Across five models, probes average accuracy and move in the expected direction on of 226 context-flip pairs. Our masking achieves the highest Completeness on every model ( versus for CI-CoT), while averaging 178 generated tokens compared with 475 for CI-CoT. These results show that contextual privacy signals in LLM representations can support targeted and efficient disclosure control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.