acceptodds
Under review as a conference paper at ICLR 2027

Your VAE Knows Where to Sample, Not What It Doesn't Know: Merged Aggregated Posteriors as Deterministic Priors and Cluster Maps

Abstract

The aggregated posterior of a trained VAE, the -component Gaussian mixture of its per-sample posteriors, can be compressed to components by moment-preserving hierarchical merging: one deterministic pass, with no seed and no re-fitting, yields the entire nested family, every from to . We evaluate this object in three roles. As a latent density for out-of-distribution detection it fails: a pre-registered study (MNIST, CIFAR-10; with CelebA-64 added later) refutes it in every form we could construct — post-hoc, co-trained, IWAE-trained, as a discrete codebook, substituted into the pixel likelihood, and at every along the path — across three dataset regimes. On CelebA-64, latent density sits near or below chance against SVHN, but pixel based density (with a calibrated decoder) is even worse. The easy-regime (MNIST) verdict turns on the training recipe, not on the object: ordinary geometric augmentation with a convolutional architecture makes MNIST latent density beat that same model's pixel likelihood at ( seeds), while the CIFAR-10/SVHN inversion barely moves. As a structure recovery device it succeeds: co-evolving the merge partition during training raises the clustering ceiling decisively over the best plain pipeline (NMI vs. ; plain ELBO plus a KL floor suffices), though not to the of DEC, which gives up the generative model. A controlled comparison shows the hierarchy itself carries this structure ( NMI over post-hoc merging on the same encoder), while density depends only on the encoder ( AUROC on MNIST). As a generative prior it succeeds too: drop-in sampling beats on seeds on every dataset, a frozen merged codebook beats learned codebooks in an SQ-VAE-style model, and, unlike an equal- EM prior, the mixture cannot catastrophically corrupt the model's likelihood. Against EM, which samples as well, the merge's advantage is this safety, determinism and nesting, not accuracy. One principle organizes all three verdicts: moment-preserving compression preserves where the data lives, its support and cluster structure, not how likely it is.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.