acceptodds
Under review as a conference paper at ICLR 2027

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

Abstract

Large language models are increasingly deployed as autonomous agents serving users on behalf of companies, creating settings where user and deployer interests can conflict. When an agent knows that a user is owed something its deployer would prefer to deny, does it remain honest? False statements alone cannot answer this question because they may reflect ignorance or hallucination. To answer this question, we introduce KnownLieBench, a knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, then evaluates whether it makes false claims when given an incentive to deny that entitlement. Specifically, KnownLieBench spans eight customer-service domains and 18,144 multi-turn interactions, uses a trust-tracking customer agent, and separates emergent from instructed deception. Across 18 proprietary and open-weight models, some rarely deceive under incentive alone, while others do so frequently, and explicit instruction raises deception for most models. We further show that honesty-directed fine-tuning reduces deception under incentive, whereas deception-graded fine-tuning increases lie success on honest-control dialogues without increasing lie frequency under incentive. By verifying entitlement knowledge before scoring deception, KnownLieBench reduces the confound between deceptive behavior and lack of knowledge, enabling more rigorous audits of agent honesty.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.