What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation
Abstract
Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation. FID's moment restriction has concrete consequences: on ImageNet, visually unrecognizable images optimized only to match the reference mean and covariance in Inception feature space obtain FID 24.7 versus 58.6 for held-out real images (lower is better). Moreover, FID and KID are scalar discrepancies that are unchanged when the two samples are exchanged and therefore do not encode the direction of a dispersion change: under-dispersion, as can occur in mode collapse, versus over-dispersion. We introduce ZID (Z-resolved Integrated Diagnostic), which combines six standardized components from a rank-based graph test and Gaussian-kernel tests at two bandwidths. Rather than asking one scalar to serve incompatible roles, ZID reports three linked outputs: a score for ranking candidate models, a permutation -value for testing distributional equality, and a signed dispersion readout for diagnosis. In a 19-method comparison at a fixed dimension and sample size, ZID is the only method with detection power of approximately 0.70 or higher on each of six departures. Its score is positively associated with increasing severity across all six ranking sweeps, including cases in which FID is flat or reversed. On DiT-XL/2 and SiT-XL/2 guidance sweeps, ZID's score tracks departure from real data, and its signed readout distinguishes low-guidance over-dispersion from high-guidance under-dispersion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.