RubricEdit-Bench: Benchmarking Complex Infographic Editing with Fine-Grained Diagnostic Evaluation
Abstract
Instruction-guided image editing models have advanced rapidly, yet their capabilities on complex infographic editing remain insufficiently evaluated. Infographics tightly couple text, data, layout, and structural relations, while existing benchmarks provide limited coverage of such structured edits and holistic evaluation protocols offer little diagnostic information about specific failures. We introduce RubricEdit-Bench, a benchmark comprising 1,483 human-reviewed cases across four categories and 15 fine-grained tasks, with attributes characterizing target multiplicity, reasoning demand, and source-image complexity. RubricEdit-Bench further adopts a diagnostic evaluation protocol that decomposes each edit into typed atomic criteria and performs source-grounded, localized verification, enabling failures to be identified at the criterion level. We evaluate 15 frontier open- and closed-source editing models, with Overall scores ranging from 1.88 to 4.89 on a five-point scale. The best-performing model achieves exact success on only 85.2% of cases, indicating that complete task fulfillment remains challenging even for frontier models. Edits requiring complex reasoning show a 0.48-point performance drop compared with edits requiring little or no reasoning. Coordinated layout, data, and structural edits also remain particularly challenging. RubricEdit-Bench provides a systematic and diagnostic testbed for identifying these remaining capability gaps and guiding progress toward reliable, structure-aware image editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.