Quadratic Advantage Response Policy Optimization
Abstract
Policy optimization methods such as trust region policy optimization (TRPO) and proximal policy optimization (PPO) improve training stability through policy change constraints. However, their surrogate objectives mainly regulate policy update magnitudes, without explicitly characterizing the relationship between advantage magnitude and policy response. We derive an analytical mapping from sample-wise quadratic curvature to advantage-conditioned responses under a general smooth quadratic probability-ratio surrogate objective. Based on this analysis, we propose Quadratic Advantage Response Policy Optimization (QARPO). QARPO assigns quadratic curvature according to relative advantage magnitude, suppressing response growth under extreme advantages. It preserves response ordering across different advantage magnitudes while retaining a smooth first-order optimization form. Theoretical analysis shows that QARPO preserves first-order gradient consistency at the old-policy point and characterizes the effect of policy normalization on the resulting scalar targets. Meanwhile, its advantage-dependent curvature reshapes the local second-order optimization structure. The quadratic restoration mechanism further provides analytical control of the PPO-Clip clipping gap, establishing an ordering between the smooth quadratic and clipped surrogates. Experiments on continuous-control and discrete-control benchmarks show that QARPO effectively regulates advantage-conditioned responses through advantage-dependent curvature design. Compared with representative policy optimization methods, QARPO achieves response differentiation and competitive performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.