acceptodds
Under review as a conference paper at ICLR 2027

Standard Evaluation Metrics for Generative Image Models Overlook Hallucinations

Abstract

This paper reveals an important problem with the standard evaluation metrics used for image generation: they cannot distinguish between two kinds of flawed samples, poor inliers and hallucinations. Poor inliers preserve the semantic properties that define the target distribution but are rendered imperfectly, with defects such as blur or warped edges. Hallucinations violate those properties, while often remaining visually plausible. The distinction is semantic rather than visual, and it significantly affects how the model should be improved. We show that the scores of the standard metrics are consistent with many combinations of the two failures, so a model that hallucinates more but produces better inliers can score the same as, or even higher than, one that hallucinates less. Having identified the problem, we analyse what determines how strongly these metrics respond to hallucinations relative to inlier quality. Experiments in controlled settings and on samples from trained generative models show that hallucinations largely go unnoticed, because the feature extractors the metrics rely on often embed them close to inliers. Even when hallucinations are embedded far from inliers, averaging over all samples hides them, while a statistic that reports only the most atypical samples recovers them and ranks models differently from the standard metrics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.