MATE: Solving Contextual Markov Decision Processes with Memory of Accumulated Transition Embeddings
Abstract
We propose MATE, a simple yet effective memory architecture for Contextual Markov Decision Processes (CMDPs), a family of MDPs parameterized by an unobserved context. In CMDPs, an optimal agent adapts online by updating its posterior belief over contexts with every new transition. MATE replaces this intractable update with a memory that additively accumulates transition embeddings, leveraging the posterior's permutation invariance to retain provably sufficient expressiveness. This design avoids both the growing per-step rollout cost of Transformers and the backpropagation through time of recurrent networks. We further propose STORE, a training algorithm that exploits the additive structure of MATE to train on a random subset of transitions per episode while reusing stored embeddings for the rest. STORE keeps the low cost of truncated-window training, yet it places no limit on the memory horizon and matches full-episode training in expectation. Across diverse benchmarks, MATE matches the performance of standard sequence-model baselines while being substantially faster to compute. Replacing truncated-window training with STORE lets MATE learn dependencies longer than the subset size and yields further performance gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.