Beyond Fixed Learning Rates: Stochastic Armijo Line Search for On-Policy Policy Optimization
Abstract
Most reinforcement learning algorithms remain far from general purpose: they must be re-tuned for each new problem. The step size is a particularly challenging hyperparameter. Good step sizes depend on many interacting algorithmic components such as function approximators, objectives, and optimizers. Even the environment and experimental procedure can influence which step sizes work. Although the community has developed good default step sizes for many benchmark suites, the inherent dependence of the step size on the environment and experimental process often necessitates re-tuning on new problems or experiment setups. The optimization community has long relied on line search as a robust method of step size selection, and stochastic line searches are gaining popularity in supervised learning. Yet, a systematic exploration of line search strategies in reinforcement learning is lacking. In this paper, we present such a systematic empirical study, developing a line search method for on-policy policy optimization with PPO, which we call ***A**rgmin **A**rmijo **A**dam* (Adam3). We benchmark Adam3 across multiple popular benchmark suites and provide ablations to show failures in basic line search approaches and the utility of the additions in Adam3. Overall, we find Adam3 is easy to implement, does not increase wall-clock time and removes the step size as a hyperparameter, incurring little cost when the defaults work but performing well when they fail.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.