How Many Images Are Enough to Compare Diversity? Adaptive Sample Sizes for Text-to-Image Evaluation
Abstract
Diversity comparisons between text-to-image generators are often computed from a fixed number of images per prompt, yet there is no principled rule for choosing . We study this question for the Vendi score, whose value keeps changing as more images are generated per prompt. However, a comparison between two generators need not be equally unstable: because the two scores tend to change by similar factors, their ratio can stabilize while each score is still changing. The number of images required for this stabilization differs widely across guidance settings, diversity interventions and model architectures. To address these issues, we introduce DivShift, a prospective procedure that determines the required number of images separately for each Vendi-based comparison. Given a user-specified reference size and a tolerance , DivShift tracks the observed Vendi log-ratio, predicts how much the ratio may still change before , accounts for variation across prompts, and either stops or requests additional images. At and , the five-panel macro-average shows that DivShift saves of generations over evaluation at , while only of stopping decisions remain outside the target tolerance. DivShift therefore provides a practical protocol for determining when a Vendi-based generator comparison has stabilized sufficiently to stop sampling, avoiding both premature stopping and unnecessary generations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.