acceptodds
Under review as a conference paper at ICLR 2027

WeGenBench: A Multidimensional Diagnostic Benchmark towards Text-to-Image Model Optimization

Abstract

Recent text-to-image models achieve impressive visual fidelity, yet existing benchmarks often reduce their behavior to coarse aggregate scores and provide limited guidance for targeted improvement. We introduce WeGenBench, a bilingual, multidimensional benchmark for diagnosing text-to-image generation capabilities. It contains 4,000 prompts across general image generation and visual text rendering, evenly balanced between Chinese and English. In addition to scenario labels, each prompt is annotated with fine-grained capability tags that capture challenges such as spatial relations, attribute binding, complex actions, negation, typography, and bilingual rendering. Aggregating results over the intersection of scenarios and tags reveals localized failure patterns and suggests hypotheses for targeted data curation and post-training. We further develop interpretable vision-language-model (VLM)-based evaluation protocols for semantic alignment and aesthetic quality, together with OCR- and VLM-based metrics for visual text rendering. These protocols provide structured rationales and dimension-level diagnostics and show stronger or comparable agreement with human judgments than the evaluated baselines on our validation set. Finally, we benchmark a broad collection of state-of-the-art systems, revealing persistent compositional, linguistic, and typographic bottlenecks that can be obscured by aggregate evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.