More HASTE, More Speed: Adaptive Sequential Testing for Efficient Text-to-Image Diagnostic Evaluation
Abstract
Diagnostic benchmarks test whether text-to-image (T2I) models satisfy conditions such as object presence, attribute binding, and spatial relations across repeated generations. They typically allocate identical image budgets to every prompt, although obvious successes and failures need less evidence than conditions near the validity threshold. We introduce HASTE (Hierarchical Adaptive Sequential Testing for Evaluation), which varies the sample count without omitting any benchmark condition. Conditions are organized by semantic prerequisites: for example, "a car is present" must hold before "the car is red" can hold. HASTE evaluates prerequisite conditions first and uses their results to inform priors for more specific conditions. It then updates each condition's posterior with binary vision-language model judgments and stops when there is sufficient evidence that its success probability lies above or below the benchmark's validity threshold. A minimum sample requirement limits premature stopping, while conditions uncertain at the sample cap use the fixed-budget decision rule. Across six T2I models, HASTE achieves average speedups of 2.60× on DSG and 4.93× on FailureAtlas, with failure-set Jaccard agreement above 0.97 relative to full-budget evaluation. At matched average image budgets, HASTE attains 0.9806 Jaccard agreement on DSG versus 0.9220 for uniform sampling, and 0.9794 on FailureAtlas versus 0.6558 for uniform sampling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.