ARGO: Asymmetric Replay Grounding with Optimistic Planning
Abstract
Model-based reinforcement learning can improve sample efficiency by reusing learned dynamics, but policy optimization can exploit inaccurate predictions for actions poorly represented in real experience. We introduce ARGO, an online framework that combines replay-grounded actor learning with optimistic data collection. Its central component, support-regularized local policy improvement (SRLPI), derives an auxiliary actor target from a KL-regularized objective at replay states. A learned replay-action density anchors the target, while ensemble disagreement guides conservative action ranking and controls the strength of regularization toward that density. The actor learns from this target alongside its base objective. During collection, an actor-centered planner combines predicted return with ensemble disagreement to select among direct actor samples and local perturbations. Executed transitions enter replay and inform subsequent model and actor updates. Experiments on visual continuous-control tasks show that ARGO improves sample efficiency over strong world-model baselines, with ablations examining the roles of collection and actor learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.