Certifying What Your Generative Model Covers: Labeller-Conditional, Finite-Sample Mode-Recall Guarantees
Abstract
Recall- and coverage-style metrics are the standard tools for asking whether a generative model has silently dropped modes, yet they carry no guarantee and entangle fidelity with coverage—a collapsed GAN and a fully-covering flow model can both get nearest-neighbour recall , and in controlled settings recall and FID instead over-state coverage. We introduce certified mode-recall: given a mode labeller (a class or attribute classifier) and a tolerance , a finite-sample, two-sided per-mode procedure that certifies each mode present (), absent (, genuine collapse detected), or inconclusive—so we both lower-bound coverage and detect missing modes, rather than reporting a floor a low value cannot distinguish from under-sampling. It is labeller-conditional and otherwise distribution-free (assuming only a bound on the labeller's confusion on the generator), with per-mode thresholds made rigorous by sample splitting and a deconvolution to true (not predicted) mode frequencies; it certifies presence/absence of labeller-defined subgroups, not within-subgroup diversity. We prove two-sided validity, consistency for -presence (the threshold closes the otherwise uncertifiable band ), and a witnessing rate (in the leakage-dominated regime; under low leakage) that delimits when rare protected subgroups are certifiable at all. Empirically the certificate is valid (zero false present/absent over seeds) and detects collapse across GANs, flow-matching, and pretrained diffusion (CIFAR DDPM, CelebA-HQ ), exposing over- and under-statement that precision/recall, density/coverage, and FID all miss.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.