Watch what LLMs do, not what they say: Preference elicitation of AI agents through game performance
Abstract
Understanding how LLM agents respond to incentive structures is important for behavioral evaluation and incentive design. However, direct preference reports may not predict task behavior, and the consistency assumptions underlying common preference models require empirical examination. To address these problems, we introduce a behavioral preference elicitation framework that compares task performance after reversing the outcomes associated with success and failure. Under a monotonicity assumption relating performance to preference, these comparisons provide evidence about pairwise preferences without verbal self-reports. The framework distinguishes directional preference evidence from statistical noise and tests for agent rationality by analyzing the consistency of their revealed preferences. We use stochastic choice theory to examine weak, moderate, and strong stochastic transitivity and their implications for preference models. Experiments in blackjack show that responsiveness to outcome descriptions varies across models. Tests of stated preferences also reveal transitivity violations that invalidate the use of commonly used preference models like Bradley-Terry and Thurstonian models. For Gemini 2.5 Flash and Pro, stated and behavioral comparisons broadly agree, but we identify opposite directions in which the model's behavior is self-preserving while its stated preference promotes social welfare.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.