acceptodds
Under review as a conference paper at ICLR 2027

Deterministic Actor–Critic Without an Action-Gradient

Abstract

Deterministic policy gradient methods such as DDPG and TD3 improve the actor by ascending the action-gradient of a learned critic, and their analysis assumes that the action value is differentiable in the action. That assumption fails on agents trained on commonly used continuous control benchmarks, and whether it fails is decided by the trained agent and not by the task alone. Differentiability in the action is governed by a competition between the discount factor and the rate at which the closed loop separates nearby states. We give the threshold below which the action-gradient exists, together with the stricter one below which the action-curvature does, and show that a loop that stretches fast enough on a hyperbolic attractor leaves no gradient at almost every state of its basin. The actor ascends the gradient of an action value smoothed by the noise TD3 adds to its targets and by the resolution of the critic. The width of the smoothing sets the size of that gradient: it is at most an inverse power of the width, and in the mean square near the attractor it grows as an inverse power when the noise is reduced. Averaged over the states the actor visits, this slope is the gradient of the return, up to a positive factor and the square of the noise, which is why the method still works beyond the threshold. On agents trained on standard benchmarks we measure the separation rate and the exponent it predicts, test the prediction against direct measurements of the action value, and give a diagnostic that detects, from the simulator alone, when a trained agent lies beyond the threshold.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.