acceptodds
Under review as a conference paper at ICLR 2027

The FID Lottery: Quantifying Hidden Randomness in Generative Model Evaluation

Abstract

The Fréchet Inception Distance (FID) is the de facto arbiter of image generation, yet most papers report just a single number from a single trained model using a single sampling seed. How reproducible is that number if we retrain the model, or merely resample from it? In this paper, we treat FID as a random variable on a two-axis panel of training and generation seeds, and measure its variance directly on several hundred SiT networks trained on class-conditional ImageNet 256×256. We report surprising findings: (a) Retraining the model using the same recipe with a different seed moves FID 3.2× more (in Inception feature space) than redrawing samples from a fixed network. (b) That gap is driven by three factors: random initialisation, data ordering, and the per-step Gaussian noise of the flow-matching loss. (c) Increasing compute or model size barely tightens the spread, holding the FID coefficient of variation (CoV) inside a 1–2% band. (d) Per-cell classifier-free-guidance tuning halves the spread but perturbs which seeds rank best, and a lucky training seed reaches the same FID with up to 2× less compute than an unlucky one. (e) The lottery extends beyond ImageNet and Inception features: on a text-to-image model retrained from 10 seeds and scored under 10 evaluation seeds each, the training seed carries 76% of the GenEval and 85% of the DPG-Bench variance. Based on these findings, we recommend a new FID evaluation protocol: evaluate under per-cell optimal guidance, treat any gap between two single runs below ≈4% of the baseline FID as inconclusive (the threshold implied by the ≈1.3% CoV floor we measure for SiT/DiT-family latent flow matching), and report an error bar over several training seeds rather than a single FID number.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.