DSPG: Learning Distribution-Conditioned Policies for Heterogeneous-Agent Economies
Abstract
With richer micro data and cheaper computation, macroeconomic policy evaluation has shifted from aggregates to the distribution behind them. Heterogeneous-agent models are the standard tool for that question, yet solving them in general equilibrium means confronting an infinite-dimensional fixed point: household policies govern how the cross-sectional distribution of assets and income evolves, and that distribution feeds back into policies through equilibrium prices. Existing solvers face a trade-off between distributional information and computational cost: low-dimensional methods (Krusell-Smith, DeepHAM, StructuralRL) reduce the distribution to a few moments, while high-dimensional methods such as DEQN take the distribution itself as input but solve too slowly to be practical. We propose DSPG (Distribution-based Structural Policy Gradient), which introduces the complete cross-sectional distribution directly into structural reinforcement learning. In a Huggett economy with aggregate risk and an explicit market-clearing condition, DSPG attains a low Euler residual of 2.5×10⁻³ in 956 s, against 6.5×10⁻³ for the strongest convergent baseline solver. A U-Net is the best DSPG network, which needs only 783 parameters to beat a 480k-parameter MLP (614× smaller) at only 3.5% higher runtime cost, and whose parameter count does not grow with the asset grid, the dominant dimension of the distribution. DSPG is the state of the art on this economy once accuracy, runtime, and model size are considered together, and it is robust across all 75 economies of a sweep that varies aggregate risk, idiosyncratic risk, and risk aversion, without retuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.