Privacy Is Elicited, Not Innate: How LLM Agents Form Disclosure Decisions, and Why Leaking Less Is Not Deciding Correctly
Abstract
An agent that acts on personal information fails visibly only after the action, when the information has already leaked. Contextual Integrity (CI) names this failure as a flow of information crossing the norms of its source context. We ask whether an LLM agent represents that norm when it decides to disclose, and where the decision forms. To read the norm and the decision from the same hidden state, we introduce CI-Pairs, variants of PrivacyLens scenarios that each change the flow and keep the rest of the story, so that the norm varies within a scenario. When the model acts, the prompt does not fix the outcome, since in a quarter of PrivacyLens scenarios some samples leak and others withhold, and before the action neither a linear nor a nonlinear probe reads the norm from the residual stream beyond what a text baseline reads from the prompt. When the model is instead asked to judge the same flow, described by its attributes, a linear probe reads the norm before it answers, with an AUROC above a text baseline on the same prompt. The model can thus represent the norm, but acting does not call on it, so its privacy behavior is elicited, not innate. We find that the decision forms inside the reasoning trace, which largely settles it within its first quarter, and that under a privacy instruction, or after fine-tuning with CI theory in the training prompt, a probe reads it from the trace beyond what the text carries. Once the action is written, a text baseline on the same words largely matches the probe, so probe AUROC alone does not show a privacy representation. A privacy instruction lowers leakage on PrivacyLens and on a small out-of-distribution benchmark, whereas fine-tuning lowers it on PrivacyLens alone, in two of our three fine-tuned models, and the one of them whose trace we can probe makes the decision no more readable than the base model does. Under CI, a correct decision also lets through the flows the norm permits, which a benchmark made only of violations cannot test, so we test this on CI-Pairs. There, the interventions that cut violations also cut appropriate disclosures, and none makes the agent detectably better at telling the two apart. Privacy must therefore be elicited, and the interventions that elicit it make the agent disclose less without making it follow the norm more closely, a poor footing for agents that act on personal information.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.