acceptodds
Under review as a conference paper at ICLR 2027

TAPE: A CELLULAR AUTOMATA BENCHMARK FOR EVALUATING RULE-SHIFT GENERALIZATION IN REINFORCEMENT LEARNING

Abstract

Out-of-distribution generalization in reinforcement learning is difficult to attribute when benchmark shifts jointly alter dynamics, observations, goals, and rewards. TAPE isolates one latent non-stationarity source—rule-shift in transition dynamics—while fixing the observation-action interface. The protocol combines deterministic splits, 20-seed replication, bootstrap uncertainty quantification, and continuous endpoints for sparse-success regimes. Across baseline families, experiments show a stable ID-to-OOD degradation pattern with strong heterogeneity across stable, periodic, and chaotic rules. The same ID-OOD gap persists in an intentionally simple ID deterministic regime, indicating that current RL pipelines remain brittle under latent-law variation even after removing common perceptual confounds. We report a protocol-matched budgeted true-dynamics random-shooting reference (p_oracle ≈ 0.187), define oracle-normalized scores ON(p) = 100 * p / p_oracle, and include a feasibility regime (L = H = 16) where rule-wise solvability reaches 100%. Under this design, low absolute success in learned agents is interpreted as a stress signal for latent-mechanism adaptation, positioning TAPE as a mechanism-oriented benchmark for robustness under explicit distribution shift.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.