CrossMaze: Benchmarking Zero-Shot Structural Generalization in Offline Continuous Control
Abstract
Policies trained on offline data are typically evaluated in familiar configurations represented in the training data, leaving their ability to transfer to unseen environments uncertain. We introduce CrossMaze, a benchmark for zero-shot structural generalization in offline continuous control. CrossMaze combines M16, a fixed diagnostic suite, with SeedMap, a procedural track for scaling training-layout diversity, in both PointMaze and AntMaze. We use CrossMaze to systematically evaluate eight imitation and offline reinforcement learning (RL) methods, including behavior cloning and offline goal-conditioned reinforcement learning (GCRL) across algorithm families and training-layout scales. As a strong reference baseline for CrossMaze, we introduce PIDA, a pretrained language-conditioned behavior-cloning policy that maps serialized maze layouts, states, and goals directly to continuous control actions. On M16, PIDA achieves held-out success rate of on PointMaze and on AntMaze, compared with and for the best evaluated non-PIDA baselines, respectively. Scaling experiments reveal that broader training-layout coverage yields uneven gains across algorithms and embodiments, while PIDA remains a strong reference with fewer layouts under the stated protocol. CrossMaze provides a controlled, reproducible testbed for measuring structural generalization in offline continuous control. Source code: https://anonymous.4open.science/r/llm_offline-B54D/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.