Constraint-Enhanced Reinforcement Learning Based on Dynamic Decoupled Spherical Radial Squashing
Abstract
When deploying reinforcement learning policies to physical robots, actuator rate constraints—hard limits on how fast each joint command may change per control step—must be respected. These limits can vary substantially across joints, producing a high-dimensional box in action-increment space. Existing solver-based projections introduce runtime optimization and training–execution inconsistency, whereas single-radius spherical parameterizations inscribe an isotropic ball and increasingly under-cover the true feasible set as dimensionality and rate-limit heterogeneity grow. This paper proposes Dynamic Decoupled Spherical Radial Squashing (DD-SRad), which computes a position-adaptive effective radius independently for each actuator and tightly aligns the closure of its reachable set with the heterogeneous feasible box. For every latent action, DD-SRad deterministically satisfies the per-step rate and position constraints. Its analytic, per-dimension mapping has a diagonally positive Jacobian at finite latent values, polynomial gradient decay, and requires no runtime optimization solver. Across MuJoCo benchmarks, DD-SRad attains the highest return among the constrained methods with zero constraint violations and improves constraint-space utilization by 30%–50% over spherical baselines. High-fidelity IsaacLab experiments on Unitree H1 and G1, using specification-consistent rate limits and native task reward configurations, provide hardware-relevant simulation evidence while not constituting physical-robot validation. Code and rollout videos are available at https://anonymous.4open.science/r/DD-SRad.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.