acceptodds
Under review as a conference paper at ICLR 2027

Drawing Answers: Benchmarking Visual Reasoning in Image Generative Models

Abstract

Can image generative models solve a problem by drawing its answer? Tracing a route, completing a grid, and selecting a winning move require images that satisfy explicit rules and spatial constraints. We introduce Drawing Answers, a dataset and benchmark spanning 22 tasks in seven reasoning families. It contains 152,366 problem–answer image pairs across 26 configurations, including four paired representations of the same underlying problems. Task-specific generators, solvers, and structured annotations support supervised learning and evaluation of generated solutions. We evaluate unadapted FLUX.1 Kontext and GPT Image 2, and train task-specific FLUX LoRA adapters, evaluating six training checkpoints and reporting the highest test outcome per configuration. Gomoku peaks at 30,000 steps, where success rises from 0.0% for unadapted FLUX to 92.0%, compared with 10.0% for GPT Image 2. Performance varies substantially across tasks: accurate local edits coexist with persistent difficulty in globally constrained solutions. Paired encoding experiments further reveal different representation preferences across models. Drawing Answers provides a controlled resource for studying the capabilities, learnability, and remaining challenges of visual reasoning through image generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.