acceptodds
Under review as a conference paper at ICLR 2027

Default Kernel Distances Miss Covariance Shape at the Sample Sizes in Use

Abstract

Practitioners compare generative models by distances between embedded sample sets: the Fr\'echet family, and kernel metrics such as the Kernel Inception Distance (KID) and the Kernel Audio Distance (KAD), which were proposed to avoid its Gaussian assumptions and each ship with a fixed default kernel inherited from the two-sample testing literature. Which differences these defaults can and cannot detect on real embedding distributions has not been measured, to our knowledge, so a practitioner reading a reported value cannot know which differences it reflects; we provide that measurement and the reporting protocol that follows from it. We treat each configuration as a statistical test and inject known distortions into real embedding distributions from symbolic music, audio, and vision. Neither default responds to a trace-preserving covariance change (a rotation) of equal or larger transport magnitude than a mean shift it always detects, and the two fail differently. KAD's median-heuristic bandwidth detects location and scale: it rejects an isotropic contraction matched in transport to the mean shift in every run in every space, and against rotations it is sample-inefficient rather than insensitive, reaching full power at sample sizes that vision datasets afford and per-genre symbolic subsets do not. KID's cubic kernel detects the mean shift and little else: it misses the matched contraction in most runs, shows no recovery at the sample sizes we reach, and does not reject on a real model against its own training data once the location component is removed. The Fr\'echet distance, run through the same permutation protocol, detects that difference. Aggregated multi-kernel tests restore power against shape at nominal level, except against coverage losses constructed to leave the mean unchanged, and near-duplicate samples corrupt split-based calibration unless the construction controls for them. These measurements give configuration choice a tested basis and yield a reporting protocol that separates three questions a single scalar conflates: whether distributions differ, how they differ, and whether samples are copied.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.