Designing Associative Working Memory for Partially Observable Reinforcement Learning
Abstract
Partially observable RL creates two complementary memory demands: integrating a trajectory into a control state and preserving specific events for later content-based decisions. We study how the latter demand shapes a working-memory interface. Our task-to-memory formulation derives that interface from the relation a later decision must recover, the retrieval cue, and the event completing the content. For transition binding, this process yields Associative Transition Memory (ATM), a slow–fast architecture coupling a recurrent controller to a return-trained associative matrix. ATM treats each completed transition as a memory token: controller context and action form its address, while the resulting observation, outcome, and termination form its content. Two controlled delayed-association tasks test transition-aligned routing efficiency and return-trained selective writing and recall. Explicit transition routing reaches a fixed competence criterion sooner in PPO-update count than a parameter-matched implicit interface. At the matched-budget endpoint, ATM achieves higher five-seed rolling-20 training return under both association loads. With four independent associations per episode, it improves over tuned Transition-FWP by 36.2% ( vs. ) and leads on all five paired seeds. Separately trained no-read and no-write controls remain near the memoryless reference; phase traces show selective updates. These advantages support the task-to-memory view that RL associative-memory architectures should be specified from task-specific information requirements along five dimensions: memory role, address, content, update, and the learning-and-validation contract.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.