Planning with Transformers: Chain of Computation and Structured Context Windows
Abstract
Large Language Models (LLMs) have had a remarkable impact across many areas of machine learning. However, recent studies have shown that they struggle to reliably solve planning problems, particularly when compared with classical planners and planning frameworks such as PDDL. Even with recent advances in reasoning-oriented models, the extent to which LLMs can genuinely perform planning remains an open question. At the same time, theoretical results have shown that transformers, the core architecture underlying modern LLMs, are Turing-complete. In this work, we investigate this apparent gap between the theoretical computational power of LLMs and their empirical planning performance. We propose (), a computational architecture that places a transformer-based LM inside an iterative loop and leverages its strength as a pattern-matching system rather than requiring it to generate an entire plan in a single pass. uses a Structured Context Window (SCW) as external memory, allowing the model to access only the portion of the accumulated context required at each planning step and thereby maintain an approximately constant input size. Within this architecture, the LM learns the planning policy and world model and performs the arithmetic operations required during planning. We show that relatively small LMs trained from scratch can learn planning procedures from a small number of training instances and generalize to unseen problems, achieving success rates above 99% on BlocksWorld and the Pancake puzzle. We further show that can be formulated as a pushdown automaton (PDA), eliminating the need for explicit pointer operations. This formulation can solve Tower of Hanoi (TOH) instances with up to 20 disks and enables the planner to execute the depth-first search (DFS) algorithm in perfect maze environments. Theoretically, we characterize how the accumulation of irrelevant context increases planning failure probability under the standard full-context setup and how the SCW mitigates this effect.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.