Compute Efficient Deep Reinforcement Learning via Normalized GLU Block
Abstract
Massively parallel simulators in robotics and game environments have shifted reinforcement learning (RL) into a wall-clock-constrained regime, where agents must achieve strong performance within limited compute budgets. This creates a fundamental trade-off in network design: larger models provide greater expressive capacity but incur higher per-step computational cost, whereas smaller models train faster but often saturate early. Addressing this challenge requires architectures that achieve higher representational capacity per parameter. We revisit Gated Linear Units (GLUs), which have shown strong parameter efficiency in vision and language models, and investigate their effectiveness in RL. We find that naively applying GLUs to RL leads to unstable training because their multiplicative interactions amplify activation variance, causing feature norm explosion. To overcome this limitation, we propose N-GLU, a simple architectural modification that combines GLUs with pre-activation layer normalization. N-GLU stabilizes feature magnitudes while preserving the expressive benefits of multiplicative gating. We provide both theoretical motivation and empirical evidence that N-GLU produces higher-rank feature representations than standard ReLU blocks, thereby alleviating plasticity loss, a common failure mode in RL. Across a diverse set of benchmarks, including Atari 57, HumanoidBench, and MuJoCo Playground, replacing ReLU blocks with N-GLU in PQN and FastTD3 consistently improves compute efficiency over standard MLP-based architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.