acceptodds
Under review as a conference paper at ICLR 2027

Routing-Aware Masked Decision Transformer for Task-Agnostic Continual Offline Reinforcement Learning

Abstract

Continual offline reinforcement learning (CORL) offers an efficient approach to building autonomous decision-making agents from static datasets without additional environment interactions during training. However, substantial differences in environment dynamics across tasks make it challenging to preserve previously acquired knowledge while effectively transferring it to subsequent tasks. Many existing structure-based methods manage knowledge through coarse-grained module composition, while their reliance on explicit task identities at inference limits their applicability when such information is unavailable. To address these challenges, we propose the Routing-Aware Masked Decision Transformer (RMDT), which integrates trajectory-based routing with parameter-level masking for task-agnostic CORL. Specifically, RMDT allocates subnetworks using learnable binary masks to preserve and reuse historical knowledge, while applying capacity regularization to discourage unnecessary parameter expansion. Gated low-rank expansion supplements model capacity on demand to accommodate task-specific learning requirements. To handle scenarios without explicit semantic task labels, RMDT introduces a lightweight task router during training that combines classifier predictions with Gaussian prototype matching. During inference, the router identifies the current task using multi-step probability aggregation and confidence-based locking, thereby activating the corresponding mask configuration to enable hierarchical knowledge selection. Experiments on Offline Continual World show that RMDT outperforms existing baselines, demonstrating better knowledge reuse and transfer capabilities. Code is available at here.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.