On Scaling Policy Networks in Policy-Gradient Reinforcement Learning
Abstract
Scaling network capacity has produced reliable gains in supervised learning, but scaling policy networks in reinforcement learning remains far less predictable. We study how policy depth interacts with optimization and interaction budget using a simple on-policy policy-gradient learner with Monte Carlo reward-to-go and a learned value baseline, alongside supporting bootstrapped variants. Across continuous-control experiments, we find that optimization conditions do not remain comparable as depth changes. Performance is highly sensitive to learning rate, while the region associated with strong performance shifts systematically with policy depth. As a result, transferring a learning rate selected for a shallow policy can confound additional capacity with optimization mismatch. Selecting learning rates separately by depth substantially changes the apparent scaling curve and reduces the measured penalty for 32-layer policies by about two fifths in our primary experiment. Reversing the transfer direction can also change the apparent depth ordering. Correcting this mismatch does not make deeper policies uniformly better. At moderate interaction budgets, deeper policies generally remain worse than shallow ones, while longer training produces task-dependent gains rather than a universal crossover. Our results show that policy-network scaling is jointly shaped by architectural depth, optimization protocol, and interaction budget. Shallow-rate transfer and per-depth selection answer different scientific questions and can yield materially different scaling conclusions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.