Can Heads Explore Better Than One in Online Reinforcement Learning?
Abstract
In reinforcement learning, generative policies can represent diverse behaviors, yet effective exploration remains a challenge in online reinforcement learning. Ensembles offer a natural approach, but training policy heads against a shared objective can reduce the diversity needed for exploration. We introduce Ensemble Sampling Actor-Critic (ESAC), a framework that combines generative policy populations with critic guided training and action selection. Circular critic pairing supplies head specific value weightings for diffusion policy regression, while pooled proposals allow the agent to explore actions from the full population. For fixed regression objectives, we prove contraction in output space under a shared objective and characterize when critic dependent targets induce distinct update operators. We further introduce Time Indexed Disagreement Exploration (TIDE), which uses policy disagreement across denoising times to guide action selection and calibrates its bonus through correlation with critic disagreement. Experiments with diffusion and flow matching policies on eight dense reward continuous control benchmarks over five seeds show improved returns over single policy controls on several locomotion tasks. Controlled comparisons and behavioral analyses examine the contribution of the policy population, while ablations establish the task-dependent benefit of calibrated disagreement. These results demonstrate the potential of generative policy ensembles for directed exploration and characterize the trade-off between improved returns and additional computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.