From Tangent Geometry to Finite-Step Curvature in Safe Policy Optimization
Abstract
Constrained trust-region methods choose updates using first-order tangent models, but evaluate finite policy steps with nonlinear reward and cost surrogates. We characterize exactly what the tangent model determines and how curvature first alters that decision at finite step sizes. With linearized constraints, the Fisher trust region's complete reward–constraint score image is preserved by , the unique smallest linear subspace with this property. Reward–Constraint Subspace (RCS) uses this tangent-exact domain to reoptimize projected reward and constraint curvature, then applies empirical-KL correction and nonlinear same-batch checks before committing an update. We also quantify the approximation from restricting updates to : along a regular small-radius branch, the full-space and restricted solutions share the same tangent term and first differ by an Fisher-orthogonal displacement governed by cross-domain active-Lagrangian curvature. The resulting reward-surrogate gap is . Under weaker Fisher-relative remainder bounds, projection provides an finite-surrogate comparator for arbitrary . On Safety-Gymnasium, matched curvature diagnostics show improved finite-step model fidelity and update reorientation, while an independent nonlinear full-space reference recovers the predicted leakage magnitude, scaling, and direction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.