Gaussian Joint Embeddings For Self-Supervised Representation Learning
Abstract
Self-supervised representation learning traditionally relies on deterministic predictions or contrastive objectives, which struggle with multi-modal ambiguities and often require heuristic asymmetries to prevent dimensional collapse. We propose a shift to generative joint modeling via Gaussian Mixture Joint Embeddings (GMJE). By explicitly modeling the joint density of context and target representations, GMJE replaces black-box prediction with closed-form conditional inference, natively providing uncertainty estimation and covariance-aware geometric regularization. We identify a fundamental optimization vulnerability in empirical batch modeling, the Mahalanobis Trace Trap, and resolve it through a suite of robust parametric, topological, and adaptive architectures. Further, we show that standard contrastive learning is a degenerate, non-parametric special case of GMJE, motivating our introduction of a dynamic Sequential Monte Carlo (SMC) memory bank to replace standard uniform FIFO queues. Empirical evaluations confirm that GMJE recovers complex multi-modal structures, achieves competitive discriminative performance, and constructs a continuous latent manifold capable of high-fidelity unconditional generative sampling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.