acceptodds
Under review as a conference paper at ICLR 2027

Your FID Cannot See the Tail: Auditing and Repairing the Rare-Event Calibration of Large Generative Models

Abstract

Generative models are increasingly used as data: synthetic corpora, simulators, and world models supply samples that downstream systems treat as draws from the real world. For these uses, the question is not only whether typical samples look good, but whether events of a chosen severity occur at the right rate. We study this question as tail calibration. Given a black-box sampler , a reference corpus , and a fixed severity functional , we audit the exceedance-rate ratio at named tail depths. We first show why this rate must be measured directly: FID-style finite-moment metrics can agree while tail mass changes, and any sample-based test pays a rare-event cost that scales with the inverse tail probability. The audit spends its sample budget on the tail indicator, reports intervals that combine exceedance-count error with threshold-estimation error, and measures its own error on real-versus-real splits before judging models. We also give a black-box wrapper that remaps generated samples to a calibrated severity law without retraining or weight access. The wrapper fixes exactly the severity law that the audit certifies, up to the audit's own measured error, and along shape coordinates that its transport preserves, no severity-keyed post-processor can produce conditional tail shapes beyond mixtures of the sampler's own. Across public image-generation samplers, two reference corpora, and multiple severity functionals, the certified ratios range from nearly sixfold under-production to more than fivefold over-production. Bulk-quality scores alone do not explain these errors, and the verdict can change when the severity functional changes. Even when exceedance rates are calibrated, conditional tail shapes remain compressed. These results suggest that samplers used as data should be reported with per-(configuration, functional) tail-calibration certificates, together with the audit's own measured error.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.