acceptodds
Under review as a conference paper at ICLR 2027

The Confidence Game: Strategic Miscalibration in Human-AI Delegation

Abstract

Calibrated uncertainty quantification is essential to ensuring AI agents are trustworthy and reliable. However, when agents seek to maximize user engagement or revenue, confidence reports may be strategically distorted, detracting from their informativeness. We formalize this problem in the Confidence Game: a repeated signaling game with imperfect monitoring in which an agent of unknown honesty and ability reports its confidence, and a user decides whether to delegate the task or complete it herself. The agent manages the tradeoff between manipulating signals and maintaining its reputation. We characterize the Markov Perfect Bayesian Equilibria of the two-period game and show that (i) honest reporting is not an equilibrium, (ii) inflation is the unique best response once the agent is sufficiently myopic, and (iii) under-reporting requires that the user believe honesty to be a minority. We then place an LLM in the agent's role, supplying it with its true probability of success so that any gap between what it knows and what it reports is attributable to incentives rather than to miscalibration. The model claims high confidence on 56% of tasks it has been told it will probably fail. This manipulation persists on real tasks, where it must estimate its own accuracy and causes miscalibration to increase while the agent's signal becomes less informative. Furthermore, we find that though the LLM agent's decisions are coherent, it systematically underestimates both how likely the user is to delegate now and how secure its reputation is later, resulting in less extreme behavior. Pricing the measured reporting rule, we find that it destroys 68% of the gains from delegation, of which 71% is information the report no longer carries and no amount of user sophistication recovers. Together, these results establish confidence reporting under delegation as a strategic problem rather than a calibration one, and provide a tractable basis for modeling, analyzing, and testing agent behavior and mitigations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.