acceptodds
Under review as a conference paper at ICLR 2027

Temporally Robust Reinforcement Learning

Abstract

Reinforcement learning (RL) agents face two kinds of distribution shift at deployment: changing *physics* (mass, friction, dynamics) and changing *clock* (delays, missed observations, asynchronous control). Although robust RL addresses the first, the second remains much less addressed, with existing work mostly confined to restrictive i.i.d. delay models. In this work, we introduce decision-time-robust semi-Markov decision processes (DTR-SMDPs), where the interaction time between the agent and the environment is state-dependent and chosen adversarially from an uncertainty set. We prove that the general problem is NP-hard, and propose a rectangularity assumption under which a policy iteration scheme converges. Our central insight is an equivalence between rectangular DTR-SMDPs and a transition-robust MDP class, with a physically plausible uncertainty set, free of the impossible transitions that standard norm-ball sets admit. Within this class, resilience to bad timing *is* resilience to bad physics. This equivalence has a practical consequence: an agent can achieve dynamics generalization by training against an adversarial clock, with no simulator-side physics randomization. We develop a practical algorithm and empirically confirm both its temporal robustness and its zero-shot transfer to physical perturbations. Finally, we characterize which robust-RL baselines do and do not inherit temporal robustness, a split our theory predicts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.