acceptodds
Under review as a conference paper at ICLR 2027

Seemingly Simple Planning Problems are Computationally Challenging: The Countdown Game

Abstract

Current foundation models and agents struggle with long-term planning, yet existing benchmarks fail to adequately assess this limitation. Most existing benchmarks either focus on loosely defined tasks like travel planning or end up leveraging existing domains and problems from international planning competitions. While the former tasks are hard to formalize and verify, the latter were specifically designed to test and challenge the weaknesses of existing automated planners. The game called Countdown, in which a player forms a target number from a list of inputs through arithmetic operations, has recently been adopted as a planning benchmark in several works, yet its computational properties and instance-space structure remain poorly understood. To address this gap, we provide the first systematic study of Countdown as a planning benchmark and propose a procedure for generating challenging instances. We argue this problem meets many of the key planning benchmark desiderata including intuitive natural language description, computationally challenging (NP-complete), and a rich instance space that avoids memorization. We provide a theoretical complexity analysis, show our instance generation outperforms public benchmarks, and evaluate existing LLM-based planners on these instances. Our results show that, unlike other domains like the 24 Game (a special case of Countdown), our proposed dynamic benchmark remains extremely challenging for existing LLM-based approaches. A key finding is the discovery of two phase transitions in problem difficulty as input size grows: a natural easy-to-hard transition, and a surprising hard-to-easy transition at larger sizes. The second transition means that naively scaling input size can overestimate a system’s planning capabilities rather than stress-test them, with important implications for how Countdown-based evaluations should be designed and interpreted.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.