acceptodds
Under review as a conference paper at ICLR 2027

CrossMaze: Benchmarking Zero-Shot Structural Generalization in Offline Continuous Control

Abstract

Policies trained on offline data are typically evaluated in familiar configurations represented in the training data, leaving their ability to transfer to unseen environments uncertain. We introduce CrossMaze, a benchmark for zero-shot structural generalization in offline continuous control. CrossMaze combines M16, a fixed diagnostic suite, with SeedMap, a procedural track for scaling training-layout diversity, in both PointMaze and AntMaze. We use CrossMaze to systematically evaluate eight imitation and offline reinforcement learning (RL) methods, including behavior cloning and offline goal-conditioned reinforcement learning (GCRL) across algorithm families and training-layout scales. As a strong reference baseline for CrossMaze, we introduce PIDA, a pretrained language-conditioned behavior-cloning policy that maps serialized maze layouts, states, and goals directly to continuous control actions. On M16, PIDA achieves held-out success rate of on PointMaze and on AntMaze, compared with and for the best evaluated non-PIDA baselines, respectively. Scaling experiments reveal that broader training-layout coverage yields uneven gains across algorithms and embodiments, while PIDA remains a strong reference with fewer layouts under the stated protocol. CrossMaze provides a controlled, reproducible testbed for measuring structural generalization in offline continuous control. Source code: https://anonymous.4open.science/r/llm_offline-B54D/

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.