acceptodds
Under review as a conference paper at ICLR 2027

DesignerBench: Benchmarking Image Generation for Real-World Design Workflows

Abstract

Image generative models are increasingly used in professional design, yet evaluation for design applications has not matured into a coherent practice attuned to real design workflows. Evaluation of generative systems in design is not a question of overall output preference, but of whether the operational requirements of a given workflow are satisfied. Existing benchmarks, however, remain largely confined to text-to-image synthesis and generic editing, leaving workflow-specific failures largely unexamined. To close this gap, we introduce DesignerBench, a workflow-centric benchmark that evaluates generative models as design collaborators, spanning 2,000 curated instances across seven design task families: asset creation, element extraction, multi-source composition, professional editing, rendering, multi-view dissection, and design language extension, which together reflect the canonical progression of real design projects. Rather than reducing evaluation to a single holistic score, DesignerBench scores each output against a sample-specific, multidimensional checklist of task-specific objectives and preservation constraints, using an evidence-first vision-language judge with a four-level rubric. Our evaluation of 14 state-of-the-art models reveals substantial variation in performance across design tasks, with rankings shifting according to the operation under evaluation. These results challenge the notion of a single best generative model, underscoring the need for workflow-aware evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.