acceptodds
Under review as a conference paper at ICLR 2027

Corruption-Robust Offline Reinforcement Learning via Causal Dynamics Adherence

Abstract

Offline reinforcement learning (RL) via sequence modeling is inherently vulnerable to training data corruptions. We observe that while data can be corrupted, the causal dynamics of the environment, such as state transitions, can remain invariant. Motivated by this key observation, we propose Causal Dynamics Adherence (CDA), a mechanism that disentangles task-relevant causal information from corruptions by constraining representations to adhere to environmental causal dynamics. To mitigate the attention leakage to corrupted inputs induced by Softmax in Decision Transformer (DT), we introduce Adaptive Sparse Attention (ASA), a plug-and-play module that augments Sparsemax with layer-specific learnable temperatures to adaptively suppress such leakage. On the D4RL benchmark, CDA-DT, which integrates CDA and ASA into DT, achieves the highest scores on 47 out of 54 (87%) tasks. The code is provided in the supplementary material and will be publicly released. Training results are available via an anonymous link.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.