Drifting through Annealed Clustering
Abstract
Drifting models generate in a single forward pass, but learning coarse structure and fine detail together remains challenging. We turn soft clustering into a novel formulation of drifting that learns the data one scale at a time. Replacing the finite centroids of deterministic-annealing clustering with a neural generator yields a new likelihood objective: explain the data through a kernel-smoothed generated distribution. We characterize its optimum, showing that when exact matching is possible, the generator represents a sharpened version of the data while the kernel gives the remaining variation. Annealing the bandwidth progressively transfers this variation to the generator, turning broad clusters into increasingly detailed distributions. We derive a simple regression algorithm with expectation-maximization targets, in which generated samples compete to explain the data without explicit repulsion or manual cluster splitting, and use it to generate reliable samples on both synthetic multi-scale data (where drifting struggles) and realistic large datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.