acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Trust Radius: Objective-Aware Steps for Policy Learning

Abstract

A fixed trust radius controls policy movement without using the gap between the current loss and a certified lower bound. We connect these two scales through Chentsov–Polyak Safe Policy Optimization, which takes the smaller of a quadratic Fisher trust-region step and a smooth Polyak step. Given a certified objective lower bound and smoothness control, this is the largest step along the fixed natural-gradient direction consistent with both the lower-Taylor and movement bounds. Objective information can therefore tighten the permitted step within a fixed trust region. For constrained learning, a separate closed-form cap preserves cumulative-cost chance feasibility from a surrogate-feasible start under sub-Gaussian costs and certified cost and curvature bounds. Synthetic experiments show reduced sensitivity to oversized radii and report no chance-infeasible iterates among 20,000 tested with the safety cap. On three Safety-Gymnasium tasks, we use a running-best reward proxy for the objective bound, a Lagrangian fallback, and one shared radius. At one million environment steps, our method retains 93–97% of radius-selected trust-region policy-optimization-Lagrangian return on two navigation tasks, but 41% on pushing, with lower pushing cost; mean costs exceed the budget on all three tasks. This exposes a distinction for policy learning: controlling policy movement, scaling optimization steps, and certifying cost constraints require different information.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.