SC-PPO: Secant-Corrected Proximal Policy Optimization for Execution-Robust Navigation
Abstract
This paper investigates execution-robust navigation in confined spaces, where a learned policy must remain collision-free even when its actions are executed with small errors. While canonical reinforcement learning (RL) algorithms such as proximal policy optimization (PPO) target the maximal expected return under exact execution, we argue that a deployed robot should trade a small amount of nominal return for insensitivity to execution errors, since limited clearance and closed-loop feedback turn these errors into collisions. However, the mismatch between the parameter space where PPO updates the policy and the action space where execution errors enter makes this goal difficult to pursue: any action sensitivity in the kernel of the policy Jacobian is invisible to the parameter gradient, so a policy can be flat in parameter space yet sharp in action space. To address the challenge, we propose a model-free RL algorithm, namely secant-correction PPO (SC-PPO), which compares two gradient evaluations within PPO's frozen evaluation context and uses their difference to correct the policy update without any disturbance model. Its action-space correction operates directly where execution errors enter and yields an action-risk bound under a measurable critic-fidelity condition, and the same operator supplies a gap-based criterion that replaces KL-divergence early stopping. SC-PPO is compared with PPO and flatness-based baselines on the BARN benchmark under persistent-bias and random-noise execution faults, and achieves the highest nominal success rate and the highest average success rate under faults. We further demonstrate the trained policy on a full-size humanoid robot.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.