Uncertainty-Guided λ-Soft Value Estimation for Sample-Efficient Maximum Entropy Reinforcement Learning
Abstract
Maximum entropy actor-critic methods such as Soft Actor-Critic (SAC) are known for effective exploration and robustness, yet their learning progress hinges on a learned soft critic whose estimates are inaccurate early in training. Model-based reinforcement learning (MBRL) promises higher sample efficiency, but compounding model errors limit how far learned dynamics can be trusted. In this work, we use model-based rollouts to improve the sample efficiency of SAC. We introduce -soft value estimation, which recursively blends the critic's direct estimate with a model-based bootstrapped estimate at each step of a model rollout, and prove that under an exact model such estimates accelerate soft policy iteration, with the one-step backup of SAC as the slowest case. To set the per-step blending factor automatically, we propose uncertainty guidance, which balances critic-ensemble disagreement against dynamics-ensemble disagreement: the estimate relies on the model where the critic is uncertain and falls back to the critic for the rest of the rollout once only the model is uncertain. On four contact-rich MuJoCo tasks, -SAC is competitive in sample efficiency, wall-clock time, and robustness to tuning, as a single hyperparameter setting solves all tasks, and it reaches a final return on par with model-free methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.