From Cooperation to Winner-Takes-All: Identifiable Slot Disentanglement under Max-Pooling via Annealed Log-Sum-Exp
Abstract
Unsupervised disentangled representation learning enables interpretable concept discovery in domains ranging from images to biological systems. Recent advances establish provable identifiability, but they rely on strict structural assumptions like sparsity or the linear independence of first- and second-order derivatives in additive decoders. Violating these conditions can degrade model performance in real-world settings. For example, datasets with overlapping objects or missing colour channels can cause a single representational slot to cover multiple objects. We study max-pooled decoders to account for scenarios where concepts mask one another. We also provide a formal proof of identifiability for these generative processes requiring fewer smoothness assumptions than previous work. We propose a temperature-annealed Log-Sum-Exp (LSE) pooling framework, which approaches the identifiable max pooling with better optimisation behaviour. At low inverse temperature, this function behaves as an easily optimisable additive decoder. As the inverse temperature increases, it smoothly converges to strict max-pooling, where our guarantees hold. This mechanism dynamically routes backward gradients to the specific representational slots most responsible for the reconstruction. On synthetic datasets with occlusion, our annealed LSE decoders outperform standard additive models and recover the latents better than naive max-pooling. Our theorems also cover the softmax masks of object-centric decoders in their max limit, and on grayscale CLEVR6 such decoders separate objects more reliably than a joint decoder. Finally, we demonstrate that annealing improves the alignment of latents with annotated cell types in six of eleven single-cell transcriptomics atlases.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.