Mean-Preserving Exploration for Continuous-Time Portfolio Learning
Abstract
We study a fundamental question in continuous-time reinforcement learning: how can we introduce exploration while ensuring that the mean exploratory action exactly recovers the target optimal policy? Using continuous-time portfolio selection as a canonical setting, we investigate two exploration mechanisms. First, we formulate entropy-regularized exploratory control and solve the associated exploratory Hamilton–Jacobi–Bellman (HJB) equations under both state-independent and state-dependent temperatures. We show that, in both cases, the mean of the optimal exploratory policy coincides exactly with the target optimal policy, while the temperature controls its dispersion. Second, we develop a direct Gaussian exploration formulation and establish the same mean-preserving property for both state-independent and state-dependent variances. Our results reveal a common principle behind these formulations: exploration can be designed through policy dispersion while preserving the optimal exploitation decision through its mean. This provides a tractable framework for separating exploration from exploitation in continuous-time reinforcement learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.