acceptodds
Under review as a conference paper at ICLR 2027

Successive-Refinement Latent Variable Generative Models

Abstract

Many generative models construct samples progressively, through sequences of latent variables, denoising steps, or iterative refinements. We study when this process forms a meaningful hierarchy, in which each stage reveals new information, improves the current reconstruction, and enables coarse-to-fine control. Our focus is on hierarchical latent variable models (HLVMs), where multiple latent variables concur to build a coarse-to-fine generation pipeline. We analyze them through *successive refinement*, the classical theory of layered lossy coding, which asks whether every partial representation achieves the smallest reconstruction error permitted by the information it carries about the source. This view yields two core principles: I) control the cumulative information rate available at each stage; II) supervise every intermediate reconstruction. We design a novel objective for hierarchical models following these principles, and show it recovers the best model-realizable hierarchy at the achieved information rates. We then analyze two canonical HLVMs under the same criteria. We find the ELBO of hierarchical variational autoencoders (HVAEs) does not incentivize hierarchical representations, as it constrains only the total rate and final reconstruction, whereas denoising diffusion probabilistic models (DDPMs) realize a special case in which the encoder is fixed, and rate–distortion optimal only for Gaussian sources under MSE. We instantiate the proposed objective in an HVAE baseline, the Successive Refinement HVAE (SR-HVAE). On analytically tractable sources it matches every prediction of the SR theory. On images it makes every latent active, enables reaching rate regions the standard ELBO never targets, and improves overall generative ability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.