GEODE: Geometry-Controlled Decoder Ensembles for Robust Offline Meta-Reinforcement Learning
Abstract
Offline meta-reinforcement learning (OMRL) combines offline RL’s ability to learn from fixed datasets with meta-RL’s capacity to rapidly adapt to unseen tasks, offering a promising framework for safe, interaction-efficient, and generalizable decision-making. However, its performance can deteriorate under context and task-distribution shifts, as errors in inferred task representations are propagated through the predictive models used to learn them. We identify an overlooked source of this brittleness that arises when decoder ensembles produce diverse predictions yet respond redundantly or overly sensitively to changes in the task representation. We introduce GEODE (GEOmetry-controlled Decoder Ensembles), a framework that explicitly shapes the local geometry of predictive supervision. GEODE combines Adaptive Diversity Control to calibrate prediction dispersion, Spectral Functional Diversity to promote complementary local response directions, and Latent-Space Robustness to limit sensitivity to task-inference errors. Predictive reconstruction anchors these objectives to observed rewards and dynamics, yielding task representations that are informative, diverse, and stable without altering the test-time adaptation procedure. Our theoretical analysis provides guarantees for dispersion transfer under context shift, conditioning of local decoder responses, and robustness to task-representation errors. Extensive experiments on six MuJoCo locomotion and Meta-World manipulation benchmarks establish state-of-the-art overall performance across in-distribution and out-of-distribution adaptation, different offline dataset quality regimes, and progressively shifted task distributions under both zero-shot and one-shot evaluation. These results establish predictive geometry as a powerful design principle for robust and generalizable OMRL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.