Proximal Behavior Reshaping
Abstract
Policy optimization in continuous control commonly uses the same parametric distribution both as the object of optimization and as the behavior distribution for environment interaction. This coupling can be restrictive when effective exploration requires richer behavioral structure, while directly using learned action values for policy improvement can make learning sensitive to critic error. We introduce Proximal Behavior Reshaping (PBR), which separates the distribution directly optimized by PPO from the distribution executed for interaction. PBR retains a Gaussian policy as the optimization anchor and learns a residual flow that reshapes its sampled actions into a richer composed policy. Learned action values guide local, policy supported reshaping without directly determining the Gaussian policy update, and are trained using the current rollout together with a short buffer of recent policy data. On a goal reaching task with wall avoidance, PBR produces coherent obstacle bypassing behavior while making effective use of locally useful but imperfect action value estimates. Across six high dimensional Isaac Lab tasks spanning locomotion and manipulation, PBR exceeds PPO across the benchmark while providing strong and consistent performance relative to more direct guided and flow based alternatives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.