Distractor-Robust Reinforcement Learning via Variational Bisimulation
Abstract
Model-based reinforcement learning (MBRL) promises data efficiency and generalization, but typical reconstruction-based objectives encourage models to waste representational capacity on task-irrelevant distractors. We introduce VIBES (Variational Inference for Bisimulation-based Encoded States), a new objective that replaces pixel reconstruction with an adversarial term, which enforces that latent states suffice to predict both rewards and the next latent state. We show theoretically that, under mild assumptions, global optima of this objective correspond to encoders that induce bisimulation relations, ensuring that latent states capture task-relevant information while discarding irrelevant variation. Our method serves as a drop-in replacement for Dreamer’s model-learning component, achieves state-of-the-art performance on the Distracting Control Suite, and extends to visually complex manipulation in Realistic ManiSkill. Unlike prior approaches, it does not rely on image-specific augmentations and applies equally well to high-dimensional vector-state tasks, demonstrated on a 100-link swimmer robot. Finally, latent-space analyses (UMAP embeddings and nearest-neighbor probes) confirm that the learned representations are sensitive to task-relevant structure while invariant to distractors.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.