acceptodds
Under review as a conference paper at ICLR 2027

Your Generative Model Is Secretly a Data Dimensionality Estimator

Abstract

Local intrinsic dimension (LID) quantifies the local degrees of freedom of the data: the independent directions in which an object can vary while remaining on the data manifold. Classical estimators infer it from nearest-neighbor statistics and can require prohibitively many samples in high dimensions. Methods such as LIDL and FLIPD instead recover it from how a generative model's density changes with the scale of added Gaussian noise. Their formulas and guarantees, however, are tied to Gaussian perturbations and to a specific model output, whereas many domains train generators with different distributions. It remains open when a model already trained for its own domain yields a justified LID estimate, and at what cost. Our main result extends this density-slope criterion beyond Gaussian noise: it gives sufficient conditions on the perturbation kernel and the local data geometry under which the logarithmic scale derivative of the perturbed density recovers LID as the noise vanishes. This derivative equals the response of the posterior-mean reconstruction of the clean point to small shifts of the noisy observation, plus a correction. The correction is required at a boundary of the support but vanishes at interior points of a smooth data manifold, where the response alone recovers LID for Gaussian and admissible non-Gaussian kernels alike. Whenever a model's output determines this reconstruction, the response is computed with a few Jacobian–vector products and no additional training. Specializing the criterion gives explicit LID estimators for diffusion, independent affine flow matching including non-Gaussian sources, generalized Brownian Schrödinger bridges and scale-conditioned normalizing flows, recovering FLIPD and Gaussian flow-matching estimators as their Gaussian special cases. On LID benchmarks with Gaussian and non-Gaussian perturbations, the response alone is competitive with the full estimate and more robust when the learned correction is inaccurate; on pretrained crystal and protein models, including ones trained with non-Gaussian noise, it tracks lattice symmetry and side-chain flexibility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.