Expressivity and Statistical Trade-offs in Diffusion Policy Learning
Abstract
Diffusion-based policies have recently emerged as powerful policy parameterizations for reinforcement learning, but their theoretical properties remain largely unexplored. In this paper, we study the expressivity and finite-sample behavior of diffusion policies generated by state-conditioned diffusion processes with parameterized drifts. We show that diffusion policies with -Lipschitz drifts approximate optimal policies with error of order , where measures the complexity of the policy class, and establish a matching lower bound showing that this rate is optimal. This expressivity gain comes at a statistical cost: a larger Lipschitz budget improves approximation but makes learning more challenging, leading to a finite-sample trade-off. When the drift is parameterized by neural networks, we decompose the policy error into three components: diffusion approximation error, neural network realization error, and statistical estimation error. Balancing these terms yields a performance gap of order , where is the sample size and is the state dimension. We also discuss how additional smoothness and low-dimensional structure of the optimal policy can further sharpen these rates. Numerical experiments illustrate this approximation–estimation trade-off and suggest a practical design principle: choose the Lipschitz budget based on the available sample size, and then parameterize only the corresponding state dependence, leading to more efficient, sample-aware diffusion policy design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.