acceptodds
Under review as a conference paper at ICLR 2027

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

Abstract

Scaling off‑policy reinforcement learning (RL) through massively parallel simulation changes the data regime assumptions, under which RL algorithms are designed. Canonical stabilizers are motivated by data‑limited training, where replay buffers provide narrow state–action coverage. By contrast, massively parallel simulation offers diverse experience at high throughput, naturally challenging the canonical roles of the stabilizers in this new data regime. Through comprehensive and controlled empirical study across eight benchmark families spanning CPU‑scale locomotion, GPU‑parallel robotic simulation, dexterous manipulation, humanoid whole‑body control and we find that these stabilizers are strongly data‑regime‑dependent: parameter normalization helps under narrow replay coverage but restricts value fitting when data are abundant, clipped double‑Q can be safely relaxed in high‑throughput manipulation, and age‑biased replay weighting is broadly useful for improving learning efficiency, especially under limited network capacity. Turning this analysis into a prescription, we build WarpSAC\xspace, a regime‑aware family of off‑policy RL algorithms. WarpSAC\xspace uses Sample Weight Decay as a regime‑agnostic component for efficient exploitation and matches each regime with a prescribed variant: WarpSAC\xspace‑L (Norm ON, clipped double‑Q) for data‑limited CPU‑scale training, and WarpSAC\xspace‑A (Norm OFF, single‑Q) for data‑abundant GPU‑parallel training. appendixpurpleWarpSAC\xspace improves normalized score–step AUC over FlashSAC by across nine CPU‑scale environments and across fourteen GPU‑parallel environments; lifts UnitreeG1TransportBox‑v1 success rate from to , and gains in mean normalized wall‑time AUC on MuJoCo Playground; achieves faster sim‑to‑real deployment on Unitree G1 than FlashSAC by 36.4% in terms of wall time. These results argue that scalable off‑policy RL should adapt its stabilizers to the available data regime. Under this principle, WarpSAC\xspace advances the state of the art of scalable off‑policy RL, delivering consistent gains over FlashSAC across different data regimes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.