acceptodds
Under review as a conference paper at ICLR 2027

ScaleAC: Scale Actor-Critic by Replay Ratio

Abstract

Employing a high replay ratio, the number of agent updates per environment interaction, has emerged as a promising strategy to improve sample efficiency in reinforcement learning (RL). However, most existing efforts to scale the replay ratio stagnate at small values, and what happens inside the agent's network under a high update frequency remains largely unexamined. In this paper, we bridge this gap by scaling the replay ratio from the perspective of network internals to achieve sample-efficient RL. We identify a critical pathology whereby the dormant neuron phenomenon intensifies with increasing replay ratios, which undermines network capacity and destroys learning. To address this problem, we propose ScaleAC, a synergy of three elements based on Actor-Critic (AC) algorithms, for high-replay-ratio settings with theoretical insights. First, to tackle dormant neurons under a high update frequency, ScaleAC introduces a periodic soft network reset in RL agents. Second, to further stabilize high-replay-ratio training, ScaleAC integrates an in-target random minimization into the Q target computation. Third, to prevent overfitting at high replay ratios, ScaleAC diversifies the replay experience through data augmentations. Extensive experiments in MuJoCo and DMC demonstrate that ScaleAC scales replay ratios in both vector-based RL (up to 256) and pixel-based RL, and extends to a modern RL network architecture (i.e., SimbaV2), yielding substantial learning acceleration and performance improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.