acceptodds
Under review as a conference paper at ICLR 2027

SpireBench: Evaluating Long-Horizon Strategic Play of LLMs with Slay the Spire

Abstract

Agentic large language models now tackle complex tasks, yet strategic reasoning over long horizons is under-studied: most agentic benchmarks are short and episodic, and performance on long tasks is hard to dissect into stored knowledge, reasoning and luck. We introduce SpireBench, in which models play the complete commercial deck-builder Slay the Spire (StS), where early decisions decide long-term success. Across 20 LLMs at the game’s highest difficulty, frontier models beat heuristic baselines and the average human player but remain far behind experts, and the gap opens as the run lengthens: the strongest models win the first act as often as experts, then fall behind. What separates strong models from weak is how well they play each fight; all of them misjudge risk over a run, preserving health rather than spending it to grow stronger. A comprehensive, self-consistent renaming of every game term shows that part of their performance rests on prior knowledge of the game, and models do not improve from the memories they write across runs, which are often wrong and overly cautious. Our harness is the first to give a model the complete original game, and we compare it with existing harnesses. Overall, SpireBench offers an unsaturated long-horizon benchmark with tools to locate where and why agents fail.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.