acceptodds
Under review as a conference paper at ICLR 2027

From Evaluation to Evolution: Benchmarking Image Editing and Evolving Image Models

Abstract

Real-world image editing requires coordinating multiple transformations with preservation and reference constraints, yet aggregate scores can obscure unmet requirements. Useful evaluation must therefore reveal local failures and translate them into actionable priorities for subsequent model improvement. We introduce WeGenEditBench, a benchmark of 2,740 instances and 148 task types spanning single-image, text-intensive, and multi-reference editing. Shared, pre-generated checklists decompose each request into atomic requirements and connect scores and deductions to visual evidence across edit execution, preservation, visual quality, and task-specific correctness. To translate diagnosis into training priorities, WeGenEvoStudio combines failure analysis, retrieval-grounded data construction, supervised adaptation, and re-evaluation in a unified multi-agent framework. Among 15 evaluated configurations, GPT-Image-2 leads single-image editing with an Overall score of 4.076, while Seedream-5.0-Pro leads text-intensive editing at 4.291. Across four generation and editing benchmarks, adaptation yields task-dependent gains; on their report-selected Top-5 targets in our single-image subset, FLUX.2 [dev] improves Overall from 1.083 to 3.517, while Qwen-Image-Edit-2511 gains 13.2%. Together, our benchmark and framework provide a foundation for developing more reliable image models through interpretable evaluation and targeted, feedback-driven learning. The benchmark and code will be publicly released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.