acceptodds
Under review as a conference paper at ICLR 2027

AttriTextBench: Fine-Grained Attribute Evaluation for Unified Visual Text Generation and Editing

Abstract

Recent state-of-the-art models have advanced rapidly in both text-to-image generation and instruction-based editing. Despite this progress, improving visual text synthesis remains a challenging and evolving frontier for image generation models. However, existing benchmarks primarily focus on content accuracy or rely on holistic whole-image evaluation. The former overlooks attribute control, while the latter cannot localize failures at individual text targets. We therefore introduce , a fine-grained benchmark for attribute-level evaluation of visual text generation and editing. , AttriTextBench covers generation tasks and editing tasks such as text addition and modification across multiple rendering scenes. , we explicitly evaluate multiple attributes, such as font, size and weight. Specifically, we curate editing instructions from real user traffic and derive generation prompts from real images. Each target records its text content and applicable attributes. , we decompose each user prompt into atomic operations. For each operation, we resolve a target bounding box in which OCR measures content and a visual judge agent assesses each applicable attribute with a binary verdict. Meanwhile, masked background comparison evaluates background preservation for editing tasks. AttriTextBench contains 721 generation prompts with 10,803 text segments and 2,854 editing instructions with 3,449 operations. Experiments on 14 models show that AttriTextBench exposes fine-grained weaknesses in visual text control and provides actionable guidance for future model development.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.