acceptodds
Under review as a conference paper at ICLR 2027

Trust Regions Sell, But Who's Buying? Generalized Power-Law Surrogates for Clipped and KL-Constrained Reinforcement Learning

Abstract

Standard proximal (e.g., PPO) and trust-region (e.g., TRPO) methods stabilize reinforcement learning but optimize linear importance-ratio surrogates that do not account for the curvature of the surrogate under large likelihood-ratio displacements. We introduce Symphony, a generalized family of RL objectives derived from the \(n\)-th root likelihood ratio, \((\pi'/\pi)^1/n\). By isolating higher-order geometric contributions via a binomial expansion, Symphony unifies computationally efficient first-order clipped surrogates with KL-constrained natural gradient updates. We analyze its local geometry, showing that this power-law transformation preserves the Fisher Information Metric direction at the origin while damping ratio-driven gradient growth away from it. Evaluations across the MJX Continuous Control benchmark show substantial gains in sample efficiency and asymptotic stability, alongside competitive, environment-dependent Atari performance on visual control tasks in JAXAtari. Notably, under extended-horizon evaluations on Humanoid (two billion transitions under a shared configuration), clipped Symphony demonstrates positive gains over PPO. In a controlled 10M-step Humanoid evaluation within our common implementation, where PPO and SPO configurations were independently selected by mean screening AUC, Symphony achieves higher final returns than both PPO and SPO under each baseline's selected settings, while demonstrating decisive sample efficiency gains over SPO. Experiments across varying geometric orders and clipping ranges confirm that Symphony provides resilience across relaxed clipping boundaries (\(\epsilon=0.4\)) and ratio drift, offering a curvature-aware foundation for scalable policy optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.