acceptodds
Under review as a conference paper at ICLR 2027

When LLMs Bid Rationally, and When They Don't: A Traceable Capability Split in Sealed-Bid Auctions

Abstract

Large language models (LLMs) are increasingly proposed as autonomous economic agents, but whether they comply with rational, incentive-compatible strategies in a canonical mechanism is rarely tested against closed-form ground truth. We study this in sealed-bid auctions, testing four LLMs against second- and first-price mechanisms. We find a clean capability split: two models, Gemini 3.5 Flash-Lite and gpt-oss-120b, are statistically indistinguishable from exact truthful bidding in our primary second-price test, contrasting sharply with the well-documented human tendency to overbid relative to the risk-neutral Nash equilibrium. The other two fail in mirror-image ways: Llama 3.1 8B overbids and, as bidder groups grow, degrades to bids indistinguishable from uniform-random noise; Qwen2.5-7B underbids by a comparable magnitude. Both trace to a symmetric declarative-knowledge error: each model states one mechanism's optimal rule correctly but carries it, incorrectly, into the other, giving a mechanistic account of both failures, not just a behavioral one. This double dissociation also produces a striking aggregation artifact: naively pooling all four models in our primary truthfulness test reverses its sign and significance entirely, since the two large, opposite-signed per-model effects nearly cancel under pooling, a cautionary result for any LLM-agent benchmark that aggregates across models rather than reporting per-model effects. We also test Regret-Conditioned Prompting (RCP), an inference-time feedback loop conditioning a model on its realized counterfactual regret each round; fit separately per mechanism with an expanded replicate sample, it produces a first-price-specific correction on our weakest model that survives Holm-Bonferroni correction (), though a cluster-robust check places it at the edge of conventional significance (); a matched follow-up on our underbidding model finds no comparable effect, so RCP's benefit is model- and mechanism-specific rather than general. Code and data will be released upon publication.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.