acceptodds
Under review as a conference paper at ICLR 2027

EuroArena: Do Large Language Models Excel at Long-Horizon Planning under Competition?

Abstract

Large language models (LLMs) have demonstrated strong performance on complex tasks, such as coding and mathematics, but can they plan effectively while other agents are planning against them? Adversarial interaction introduces a distinct challenge: a plan must account not only for the consequences of one's own actions, but also for opponents who can contest resources and disrupt its execution, which is largely overlooked in current benchmarks. We introduce , a benchmark featuring two-player Eurogames that tests long-horizon planning through competition over shared resources, investment in future gains, and adaptation to opponent's actions. Across more than 400 hours of Splendor Duel gameplay against skilled human players, we find that, surprisingly, six strong LLMs struggle to match human expertise, despite demonstrating the ability to formulate concrete plans, revise them after feedback, and propose alternatives for possible future states. To investigate this gap, we introduce resource dependency graphs to trace how earlier decisions support later purchases, and leverage this with human-verified trajectory analysis. We find a recurring contrast: LLMs prioritize reaching the next subgoal, whereas humans more often take actions that support later subgoals and make better use of accumulated resources. Models also overlook adverse opponent responses and make reasoning errors when integrating multiple information sources. These findings reveal a gap between generating and revising explicit plans and making decisions that sustain progress toward long-term goals under competition.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.