CoSR: Memory-Efficient FP8 Training for Multitask Reinforcement Learning
Abstract
FP8 reduces compute and memory costs in large-scale deep learning, but bootstrapping, nonstationary targets, and closed-loop feedback between policies and data distributions make deep reinforcement learning sensitive to low-precision perturbations. High-capacity multitask critics further increase training-state memory costs. We find that deterministic writeback in FP8-resident online actor–critic training can suppress updates smaller than the quantization step, degrading long-term performance. Independent stochastic rounding (SR) preserves these updates in expectation but ignores error interactions in forward and backward matrix operations and Adam state updates, allowing locally unbiased noise to accumulate into substantial optimization perturbations. We propose CoSR, an operator-aware joint SR framework. CoSR uses forward/backward channel correlations and the coupled effects of Adam's first and second moments on update directions to coordinate correlations among rounding errors while preserving elementwise rounding probabilities and unbiasedness. It requires no per-parameter compensation buffers, reducing persistent-state memory overhead; major weight matrices and Adam states reside in FP8. Across diverse reinforcement-learning benchmarks, including DeepMind Control Suite and Meta-World, CoSR achieves performance comparable to the high-precision baseline while reducing training-state memory and peak GPU memory by 74% and 63%, respectively. These savings enable wider critics, yielding substantial multitask performance gains under the same hardware budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.