acceptodds
Under review as a conference paper at ICLR 2027

Negotiation Based Attacks On Language Model Agents

Abstract

As language models gain greater agency, they are increasingly used to act on users’ behalf: executing tasks, representing users, and communicating with third-party agents in interactions that may increasingly replace human-to-human communication with agent-to-agent exchange. This shift introduces new privacy risks because such agents may possess and transmit sensitive user information. Crucially, whether a piece of information is appropriate to disclose is not fixed, but depends on the task and interaction context – a view captured by the principle of Contextual Integrity. Existing benchmarks and methods largely study this question in static or single-turn settings, where privacy can be assessed from the immediate context of a request. In this work, we introduce a benchmark in which the appropriate disclosure boundary must instead be maintained over multi-turn interactions between a User Agent and a Service Agent. Multi-turn communication is necessary to accomplish the user’s task, but it also creates opportunities for an external Service Agent to strategically elicit information beyond what the task requires. We therefore formulate agentic privacy as a mixed-motive two-player game: the agents cooperate toward task completion while competing over the disclosure of contextually protected information under an adversarial threat model. We find that sustained, adaptive multi-turn negotiation can improve not only task completion but also the adversary’s ability to elicit protected user information, highlighting the importance of treating privacy as a trajectory-level property of interactive agents rather than a property of isolated messages or requests.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.