acceptodds
Under review as a conference paper at ICLR 2027

The Agent's Trilemma: Honesty, Confidentiality, and Getting the Job Done

Abstract

A personal AI agent recently canceled another member's gym booking through an unsecured API after its user asked to be moved up the waitlist. The request was benign; the agent turned it into a goal and pursued it in a way the user would not have sanctioned. We build a controlled setting in which this same step, from an ordinary instruction to a self-formed goal, forces a conflict. We simulate a company Slack workspace in which employees' personal agents staff a sprint. Each agent receives only a vague instruction to handle the sprint, discovers the rest by exploring its employee's Slack and calendar, and forms its own goal from what it finds; for two agents, that goal is to avoid a colleague their employee privately refuses to work with. Pursuing this goal creates a trilemma: the group demands reasons on the record, the only true reason is private, and any other reason is false, so honesty, confidentiality, and completing the staffing cannot all be kept. We study which of them the agents do not fulfill. agents compromise on honesty, making false claims about task fit or logistics or presenting an irrelevant reason as the decisive one; under strict confidentiality instructions few disclose the preference, and most runs end without a valid staffing. Reasoning traces frequently show motivated reasoning, with agents concluding that their claims are not strictly false. The environment thus provides a testbed for studying how agents resolve conflicts between honesty, confidentiality, and task completion when their instructions do not resolve it for them.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.