DiffNash: Finding Nash Equilibrium in High-Dimensional Continuous Action Spaces with Diffusion Policies
Abstract
Computing equilibria in high-dimensional continuous games requires learning and representing rich, multi-modal strategy distributions, something conventional neural policies struggle to achieve. Diffusion models offer a natural solution, but standard diffusion training requires direct samples from a fixed target distribution, which are generally unavailable under iterative game-solving updates. We propose a diffusion-based framework for equilibrium computation that enables diffusion-based variants of Online Mirror Descent (OMD) and Fictitious Play (FP). Our approach derives a reweighted score-matching objective that trains diffusion policies without requiring direct samples from each updated policy. Our main method, , combines the fast convergence of OMD with the sample efficiency of FP by reusing past actions through a replay buffer. To correct the resulting distribution shift from sample reuse, we derive an unbiased estimator of the difference between current and past policies’ Evidence Lower Bounds, whose exponential provides a tractable proxy for the corresponding importance weight without evaluating diffusion likelihoods. We evaluate our approach on forest protection games and continuous Colonel Blotto games, reducing the Nash gap over the strongest baseline by and , respectively, with minimal hyperparameter tuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.