Gromov-Wasserstein Information Bottleneck for Distance-Preserving Representation Learning
Abstract
Preserving useful relational geometry is a fundamental challenge in representation learning for generative and predictive tasks. However, task losses and distributional constraints do not directly specify how an encoder transforms relationships among samples. In this work, we study geometry-aware representation learning from an information-theoretic perspective and develop the Gromov-Wasserstein Information Bottleneck (GWIB). Drawing on the information bottleneck principle, we derive an upper bound for an empirical kernelized information quantity involving the Gromov-Monge gap. The gap measures the encoder's excess structural cost relative to an optimal Gromov-Wasserstein coupling between the same input and encoded distributions. For counterfactual regression, we couple within-group geometry with cross-group fused-GW alignment through a shared transport plan. For autoencoders, the geometric term complements native reconstruction and prior-matching objectives. Experiments on IHDP and ACIC support the coupled causal construction through treatment-effect comparisons and ablations. Image experiments compare GWIB-augmented autoencoders with their base models on FashionMNIST and CIFAR-10, assessing generation, reconstruction, and representation utility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.