Ensemble Value-Based Reinforcement Learning through the Lens of Mode Connectivity
Abstract
Ensemble methods are a common solution for improving the stability and performance of value-based reinforcement learning (RL) algorithms. However, conventional ensembles operate in the output space and require maintaining multiple models during inference, resulting in inference cost that grows linearly with the ensemble size. In this work, we study the geometric structure of the objective landscape in value-based RL through the lens of mode connectivity and analyze how trained Q-networks relate in parameter space. Our analysis reveals conditions under which weight-space ensembling becomes effective, how linear connectivity can emerge between Q-networks, and which RL-specific conditions can break this connectivity. Building on these insights, we propose Q-Soup, a weight-averaging approach that merges multiple Q-networks trained under identical or diverse configurations into a single model. Empirical results on the Atari benchmark show that Q-Soup achieves performance comparable to traditional output-space ensembles while maintaining a single model at inference time. To the best of our knowledge, this work presents the first systematic investigation of mode connectivity and weight-space ensembling in value-based RL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.