SolveEdit: Benchmarking Visual Problem Solving in Generative Models
Abstract
Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without disturbing unrelated content. However, existing benchmarks mainly evaluate perception, generation, or explicitly specified transformations, leaving goal-driven visual problem solving underexplored. To bridge this gap, we introduce \solveedit, a benchmark for visual problem solving through scene transformation. Given an image and a goal, a model must infer a valid transformation from the request, the scene, or a visually expressed rule, then execute it while preserving unrelated content. \solveedit contains 2,728 cases across 10 domains and 54 subdomains. Atomic transition contracts specify required and protected conditions, enabling \solvescore to measure completion and unintended changes without a single reference output. The strongest evaluated model achieves only 57.0% \solvescore. We further introduce \solveeditplan, a two-stage visual planner that instantiates the transition before generation. Under matched single-generation evaluation, it improves \solvescore by 9.1 points on average across three tested generators, including a gain from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.