ElectionBench: Evaluating LLM Agents through Competitive Elections
Abstract
Evaluating LLM agents requires analyzing their behavior in dynamic environments: how they accumulate information, take actions, and pursue objectives over long-horizon interactions. Yet existing benchmarks rely on static environments or single-answer tasks that lose discriminative power as model performance saturates. To address these limitations, we introduce ElectionBench, a novel competitive benchmark where LLMs campaign head-to-head to persuade a shared synthetic electorate in procedurally generated elections. Operating under incomplete information and costly action trade-offs, agents must infer various hidden information such as voter preferences, and publish persuasive articles. Across 11 environment profiles, we benchmark 15 LLM families using mirrored matches and Bradley-Terry scaling anchored to a scripted baseline. Experimental results indicate that winning models (led by Claude-Sonnet-5 and Gemma-4-31B-it) succeed through strategic restraint, taking fewer actions and preserving capital. Furthermore, closing debates overturn 24% of election outcomes, highlighting the decisive interplay between dynamic persuasion and long-horizon resource management. Together, ElectionBench provides a rigorous testbed for evaluating how autonomous agents integrate active information gathering, economic decision-making, and natural-language communication under direct competition.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.