Parameter Step-Based Exploration For Learning Deterministic Policies
Abstract
In this work, we investigate deterministic policy learning through an auxiliary objective, \(J_\sigma\), induced by step-based Gaussian perturbations of the policy parameters, where \(\sigma>0\) denotes the exploration scale. First, we establish discrepancy bounds quantifying the gap between the auxiliary and deterministic objectives under both Lipschitz and smooth Markov Decision Processes (MDPs), of order \(O(d\sigma)\) and \(O(d\sigma^2)\), respectively. Furthermore, analyzing a simultaneous actor-critic architecture under realistic Markovian sampling rather than idealized i.i.d. sampling, we prove that our approach reaches an \(\varepsilon\)-stationary point of \(J_\sigma\), up to critic approximation error, with a sample complexity of \(O(d^1.5\sigma^-3\varepsilon^-2)\) in the Lipschitz case and \(O(d\sigma^-2\varepsilon^-2)\) in the smooth case. Under additional smoothness, we also bound the gradient discrepancy between the auxiliary and deterministic objectives by \(O(d^1.5\sigma)\), thereby relating stationarity of the explored objective to stationarity of the deterministic deployment objective as the exploration scale decreases. Finally, we evaluate PSBE and a deep variant, Deep-PSBE, on standard MuJoCo continuous-control benchmarks against parameter- and action-space exploration baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.