MADE: A Memory-Augmented Distribution Expansion Approach via Sleep Consolidation for Context-Based Offline Meta-Reinforcement Learning
Abstract
Context-based offline meta-reinforcement learning acquires universal policies from fixed datasets by conditioning on task contexts, thereby enhancing agent generalization to unseen tasks. However, existing methods suffer from severe performance degradation under context shift, where test-time contexts are collected by behavior policies that differ from those seen during training. This brittleness arises because prior approaches optimize task representations solely on training data collected by fixed behavior policies, leaving agents vulnerable to out-of-distribution (OOD) failures at test time. Inspired by sleep consolidation in mammalian brains, we propose MADE, a memory-augmented framework that actively expands the task distribution in latent space during offline training, thereby enriching the diversity of context representations. MADE introduces three components: an uncertainty-aware memory bank identifying high-variance regions as expansion priorities, a counterfactual task generator synthesizing novel tasks via interpolation and conservative extrapolation, and a meta-critic filtering synthetic data through consistency, diversity, and causality assessments. Evaluations across five benchmarks show that MADE improves OOD performance by up to 13.10% compared to state-of-the-art methods and achieves up to 94.36% performance retention, effectively validating this active distribution expansion paradigm.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.