acceptodds
Under review as a conference paper at ICLR 2027

HE-FastSAC: Heterogeneous Training Objectives for Massively Parallel Off-Policy Reinforcement Learning

Abstract

Off-policy reinforcement learning has recently emerged as an alternative to on-policy methods on massively parallel simulators, where reuse of replay-buffer data improves sample efficiency. Current recipes in this regime optimize actor and critic under a single training objective, with hyperparameters tuned per task or per benchmark suite. Building on FastSAC, we present HE-FastSAC, which instead trains a population of policies under heterogeneous training objectives in a single run with minimal overhead. We partition the parallel environments into groups, each with its own target entropy and discount factor. A single actor and critic are shared across groups, and every group learns from the experience collected by all groups. We find that a policy learns faster and reaches a higher final score when trained inside this population than when trained alone under the same training objective. The gain comes both from the experience the population collects and from training a single network under multiple training objectives. Using one fixed hyperparameter set across 56 continuous-control tasks from four benchmark suites, with the same group deployed on every task, HE-FastSAC outperforms strong off-policy baselines and matches a per-task hyperparameter sweep of the same agent trained alone. Selecting a trained policy per task from the population yields a further gain with no additional training cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.