acceptodds
Under review as a conference paper at ICLR 2027

Parameter Step-Based Exploration For Learning Deterministic Policies

Abstract

In this work, we investigate deterministic policy learning through an auxiliary objective, \(J_\sigma\), induced by step-based Gaussian perturbations of the policy parameters, where \(\sigma>0\) denotes the exploration scale. First, we establish discrepancy bounds quantifying the gap between the auxiliary and deterministic objectives under both Lipschitz and smooth Markov Decision Processes (MDPs), of order \(O(d\sigma)\) and \(O(d\sigma^2)\), respectively. Furthermore, analyzing a simultaneous actor-critic architecture under realistic Markovian sampling rather than idealized i.i.d. sampling, we prove that our approach reaches an \(\varepsilon\)-stationary point of \(J_\sigma\), up to critic approximation error, with a sample complexity of \(O(d^1.5\sigma^-3\varepsilon^-2)\) in the Lipschitz case and \(O(d\sigma^-2\varepsilon^-2)\) in the smooth case. Under additional smoothness, we also bound the gradient discrepancy between the auxiliary and deterministic objectives by \(O(d^1.5\sigma)\), thereby relating stationarity of the explored objective to stationarity of the deterministic deployment objective as the exploration scale decreases. Finally, we evaluate PSBE and a deep variant, Deep-PSBE, on standard MuJoCo continuous-control benchmarks against parameter- and action-space exploration baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.