Temporal-Structure Robustness Transfer in Offline Reinforcement Learning: Uncertainty-Guided Decision Transformer
Abstract
Existing corruption-robust offline RL methods largely focus on learning reliable policies from corrupted offline datasets, while robustness transfer under changes in the temporal structure of corruption remains less understood. We formulate temporal-structure robustness transfer (TS-RT), an offline RL setting in which policies are trained under temporally independent corruption and tested under previously unseen, temporally correlated corruption. To the best of our knowledge, this work provides the first systematic investigation of TS-RT in offline RL across observation, action, and reward corruption settings. To address TS-RT, we propose Uncertainty-Guided Decision Transformer (UDT), which learns a token-level uncertainty signal to regulate both causal information routing and the training influence of unreliable targets. For controlled evaluation, we introduce Drift-Attack, which generates temporally correlated test-time perturbations through burst–recover dynamics and momentum-driven drift within a matched-channel protocol. Across nine D4RL tasks and seven representative baselines, UDT achieves the highest overall normalized score of 63.52 and ranks first on 11 out of 15 reported aggregate evaluation metrics across corruption types, corruption channels, and environment groups. We provide code and trained checkpoints to facilitate reproducibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.