AgentCIScope: Grounded Contextual Integrity Evaluation for LLM Agents
Abstract
LLM agents are becoming increasingly capable of handling end-to-end daily workspace tasks, enabled by their ability to smoothly interact with environments through tool calls. However, as LLM agents are increasingly deployed in information-rich, real-world environments where they read and write user data across email, calendar, and other applications, a central privacy question is whether they preserve contextual integrity (CI), that is, whether the information flows they produce comply with context-specific disclosure constraints. Existing agent evaluation platforms and benchmarks either only focus on model-level privacy leakage or evaluate agents across a limited set of application environments with manually constructed disclosure constraints. In this work, we introduce AgentCIScope, a unified platform with a comprehensive benchmark for evaluatingagent-level contextual integrity grounded in explicit disclosure constraints through executable, multi-step tasks across diverse, realistic application environments. AgentCIScope is constructed through a unified pipeline that transforms synthetic and real-world privacy cases into executable agent tasks, with grounded disclosure constraints defined by the intersection of three criteria: governing privacy policies, persona-based agreement, and a predefined sensitive-attribute taxonomy. Using AgentCIScope, we evaluate 12 types of agents over 936 tasks. We find that strong task-completion capability does not necessarily imply contextual-integrity preservation: single-run leakage rates range from 32.5% to 94.4%, while repeating each task five times further increases leakage by roughly 4–19 percentage points. Across agents, we find that certain sensitive attributes are leaked more often, with personal preferences and booking details reaching average leakage rates of 84.7% and 76.5%. In addition, highly sensitive attributes such as credit-card information and SSNs also exhibit average leakage rates around 17.1% and 29.4%. Among different environments, Calendar has the highest average leakage rate as an information source, while WhatsApp is ranked the highest as an information destination; leakage rates also vary substantially across governing privacy policies. Our evaluation results provide broad insights into patterns of agent privacy leakage, with AgentCIScope serving as a unified testbed for developing privacy-preserving agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.