Challenging the Manifold Hypothesis: A Geometric Diagnostic for Data and Representation Spaces
Abstract
The manifold hypothesis is widely invoked to motivate representation learning, dimensionality reduction, and geometric operations in high-dimensional data spaces. Yet it is often stated informally: data are said to lie near a low-dimensional “manifold” without specifying the geometric properties this object should satisfy or whether finite datasets exhibit them. We introduce the Geometric Manifoldness Index (GMI), a continuous diagnostic for two necessary signatures of smooth embedded manifold structure in point clouds: local intrinsic-dimension consistency and positive Federer reach. GMI is scalable to modern high-dimensional representations and separates dimension inconsistency from reach-related failures such as cusps and self-intersections. We validate GMI on synthetic manifolds, controlled singular examples, and image-transformation manifolds, then apply it to raw images, autoencoder latents, token embeddings, and single-cell transcriptomics data. Across these modalities, many data and representation spaces remain far from the smooth-manifold regime, even when low-dimensional structure is visible in projections or induced by compression. Learned embeddings become more manifold-like mainly when training objectives impose geometric regularity, and manifoldness shows no consistent positive correlation with reconstruction quality. Our results challenge common informal uses of the manifold hypothesis and motivate diagnosing manifoldness as an explicit geometric assumption rather than presuming it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.