Reinforcement Learning by Change of Measure: A BSDE Theory for Generative Policies
Abstract
Online reinforcement learning with generative policies requires translating scalar critic values into changes of the multi-step stochastic dynamics that generate actions. Existing online generative-policy methods typically focus on actor-side objectives based on value reweighting, importance sampling, or critic action gradients. We seek a unified formulation that also identifies the soft continuation value used in the Bellman update. We introduce a unified change-of-measure formulation of actor-critic updates for online RL with generative policies, in which an offline generative policy defines a reference process and critic-induced improvement is recast as a controlled change of its path law toward the desired Gibbs endpoint distribution. We prove that this change of measure induces a positive martingale whose logarithm satisfies a quadratic backward stochastic differential equation (QBSDE), identifying the soft value and the Doob-Girsanov drift correction as two components of the same probabilistic transformation. This equivalence converts Boltzmann policy improvement into a QBSDE estimation problem under the reference process. Building on this principle, we develop Value-Only BSDE-Controlled Diffusion Policies (VBCDP), which learns the value and policy drift directly from scalar critic evaluations without requiring critic action gradients. Controlled experiments recover analytic Gibbs–Doob solutions, while results on six DeepMind Control Suite tasks show consistent improvement over matched frozen references, with mean gains of -, and competitive performance against modern generative-policy baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.