META: Mamba Episodic Temporal Alignment for Stable Reinforcement Learning
Abstract
Structured State Space Models (SSMs), such as Mamba, offer an efficient mechanism for modeling long-range dependencies and are well suited to partially observable reinforcement learning (RL). However, directly applying Mamba to on-policy RL introduces two forms of temporal instability. First, we identify -drift, where the learned temporal scaling parameter , which governs the effective memory duration of the model, gradually diverges during training and destabilizes useful temporal context. Second, we observe episodic contamination, where residual state information persists across episode boundaries and violates episodic independence. To address these issues, we introduce META (Mamba Episodic Temporal Alignment), a framework for stabilizing the temporal dynamics of Mamba in episodic RL. META combines two complementary mechanisms: Episodic Forgetting, which resets temporal traces at episode boundaries to prevent cross-episode interference, and Memory Rebalancing, which regulates the sensitivity of through structured spectral reparameterization to counteract drift during training. Across eight partially observable Brax benchmarks, META improves training stability and sample efficiency over recurrent and SSM-based on-policy baselines in most settings, achieving up to more than higher final returns than the strongest prior SSM baseline, S5RL, on tasks where that baseline attains positive performance. These results highlight temporal stability as a central challenge for deploying selective state-space models in episodic reinforcement learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.