FLASH: Improving Downstream Experience Utilization in Visual Model-Based Reinforcement Learning through Task-Relevant Episodic Retrieval
Abstract
Recent advances in visual model-based reinforcement learning have continually improved control performance through higher-fidelity generative world models and visual representations, often at additional computational cost. However, our probing experiments show that lightweight latent representations can retain substantial task-relevant information even when visual reconstruction is imperfect, suggesting that representation fidelity alone may not determine downstream control performance. These findings suggest that downstream utilization of task-relevant information already available in latent representations may itself be an important bottleneck. We therefore propose Fast Latent Anchoring via Salience and Hashing (FLASH), which uses salient learning events as retrieval anchors to identify structurally similar historical contexts through a hash-indexed episodic memory. FLASH jointly reuses these contexts in Actor-Critic learning, improving downstream experience utilization without replacing the underlying encoder or world-model architectures. Across Atari 100k, FLASH improves STORM's mean and median human-normalized scores by 4.71% and 8.37%, respectively, and DRAMA's by 8.33% and 5.35%. As a plug-and-play module that can be readily integrated into existing agents, FLASH achieves these performance gains with a 5.11–17.61% increase in wall-clock training time across the evaluated configurations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.