TAPE: A CELLULAR AUTOMATA BENCHMARK FOR EVALUATING RULE-SHIFT GENERALIZATION IN REINFORCEMENT LEARNING
Abstract
Out-of-distribution generalization in reinforcement learning is difficult to attribute when benchmark shifts jointly alter dynamics, observations, goals, and rewards. TAPE isolates one latent non-stationarity source—rule-shift in transition dynamics—while fixing the observation-action interface. The protocol combines deterministic splits, 20-seed replication, bootstrap uncertainty quantification, and continuous endpoints for sparse-success regimes. Across baseline families, experiments show a stable ID-to-OOD degradation pattern with strong heterogeneity across stable, periodic, and chaotic rules. The same ID-OOD gap persists in an intentionally simple ID deterministic regime, indicating that current RL pipelines remain brittle under latent-law variation even after removing common perceptual confounds. We report a protocol-matched budgeted true-dynamics random-shooting reference (p_oracle ≈ 0.187), define oracle-normalized scores ON(p) = 100 * p / p_oracle, and include a feasibility regime (L = H = 16) where rule-wise solvability reaches 100%. Under this design, low absolute success in learned agents is interpreted as a stress signal for latent-mechanism adaptation, positioning TAPE as a mechanism-oriented benchmark for robustness under explicit distribution shift.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.