Auditing Action-Gradient Geometry in Ensemble Cost Penalties
Abstract
A common ensemble cost penalty in safe reinforcement learning adds prediction spread to the mean cost. A gradient-based actor follows the action derivative of this penalized estimate. We introduce an action-gradient audit to distinguish changes in direction from changes in signed scale and magnitude. The audit evaluates recorded behavior actions reserved from critic fitting. On matched data from three MuJoCo systems, bootstrap ensembles produce greater rotation than shared-data ensembles. The median increase exceeds 6.5° in HalfCheetah, while median parallel scales and norm ratios remain near one. Across independently seeded fits to the same data, mean gradients agree more closely than penalized gradients in every studied group. The penalized direction also has a nonpositive median advantage over the mean direction in 71 of 72 finite-step comparisons on independently fitted evaluators. These findings show that a penalty can substantially rotate the local cost gradient without yielding a direction that agrees better across fits. Evaluating both scalar estimates and their derivatives makes this distinction visible.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.