acceptodds
Under review as a conference paper at ICLR 2027

GRL: From GPU Underutilization to Faster Reinforcement Learning

Abstract

Compact networks in replay-based reinforcement learning (RL) often leave GPU capacity unused, yet increasing execution parallelism can also change the learning work needed to reach a target quality. We present GRL, a GPU execution workflow that preserves algorithm-defined learner-update structure while reorganizing collection, data delivery, and evaluation around it. Actor-side batching and cross-role overlap expose parallel work; GPU-resident handoff separates transfer completion from subsequent computation, minimizing synchronization overhead. Its learning-aware execution configurator screens candidates by throughput and compares learning quality through short trials, using a bounded search to limit selection overhead before training. Across eight environments with different complexity, GRL achieves end-to-end speedups up to over CPU-GPU asynchronous execution and over GPU-synchronous execution, measured across ten training runs sharing one configuration cost. Execution and configuration ablations show how throughput, learning progress, and selection overhead jointly shape first time-to-quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.