acceptodds
Under review as a conference paper at ICLR 2027

ProcTD: A Procedural Tower Defense Benchmark for Spatial Reasoning and Long-Horizon Planning

Abstract

Tower Defense (TD) games combine long-horizon economic planning with short-horizon spatial reasoning, yet remain largely absent from reinforcement learning benchmarks. We introduce ProcTD, a procedurally generated, Gymnasium-compatible TD environment with partial observability, configurable spatial scale, heterogeneous tower–enemy matchups, technology progression, and an optional tower-selling mechanic. Its JIT-compiled implementation exceeds 10,000 environment steps per second on a single CPU core, making large-scale experimentation feasible on commodity hardware. The raw action space grows quadratically with map size, from 1,540 discrete actions on to 6,148 on , yet only a small subset is valid at any step. We benchmark general-purpose RL methods and find that performance degrades substantially as spatial scale increases. As a feasibility baseline, we introduce Macro-Action Policy with Symbolic Placement Proximal Policy Optimization (MAPS-PPO), which combines a learned high-level policy with handcrafted spatial placement that reduces the action space by more than 80%. MAPS-PPO reaches a median of waves on with an IQR of , and again a median of on but with an IQR of , while PPO reaches and waves, respectively. ProcTD therefore provides a lightweight benchmark for studying large, structured, spatially grounded action spaces and is currently not solved by general-purpose methods. We release the environment open-source with baseline implementations and an asymmetric attacker–defender extension that allows for competitive benchmarking.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.