Inferring Latent Geometric Embeddings from Grouped Relational Data
Abstract
Graph embedding methods place nodes so that distance matches an observed graph. In many domains no such graph exists, and relatedness between entities must be inferred from repeated, grouped counts collected across heterogeneous contexts. We introduce a generative model in which relatedness is latent rather than observed. Entity positions on a product manifold determine a pairwise propensity, and an observation layer links it to overdispersed counts. Flattening the manifold and taking one observation per pair recovers a Euclidean latent-space model. We derive the variance ratio between this geometric estimate and a fit using only a pair's own counts, and show it depends on how precisely the pair's endpoints are determined. A sparse pair between well-determined entities is therefore estimated nearly as well as a data-rich one. The ratio involves the geometry only through its dimension, so it applies to flat and curved embeddings alike. The same quantities yield a combined estimator with no fitted parameters. We sample the embedding's posterior with a Riemannian extension of stochastic gradient Langevin dynamics, and because the likelihood is generative, an entity absent from training can be placed from its own evidence alone. On cancer co-mutation data, 5% of that evidence suffices to outpredict full-evidence baselines on two thirds of genes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.