acceptodds
Under review as a conference paper at ICLR 2027

MATE: Solving Contextual Markov Decision Processes with Memory of Accumulated Transition Embeddings

Abstract

We propose MATE, a simple yet effective memory architecture for Contextual Markov Decision Processes (CMDPs), a family of MDPs parameterized by an unobserved context. In CMDPs, an optimal agent adapts online by updating its posterior belief over contexts with every new transition. MATE replaces this intractable update with a memory that additively accumulates transition embeddings, leveraging the posterior's permutation invariance to retain provably sufficient expressiveness. This design avoids both the growing per-step rollout cost of Transformers and the backpropagation through time of recurrent networks. We further propose STORE, a training algorithm that exploits the additive structure of MATE to train on a random subset of transitions per episode while reusing stored embeddings for the rest. STORE keeps the low cost of truncated-window training, yet it places no limit on the memory horizon and matches full-episode training in expectation. Across diverse benchmarks, MATE matches the performance of standard sequence-model baselines while being substantially faster to compute. Replacing truncated-window training with STORE lets MATE learn dependencies longer than the subset size and yields further performance gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.