Reflection-Driven Code-with-Image Reasoning
Abstract
Many visual tasks ask for quantitative properties of an image, such as displacement, rotation angle, or color cast. Solving these tasks requires computation over image pixels rather than visual inspection alone. We study Code-with-Image, a problem-solving paradigm in which a large vision-language model designs an algorithm, implements it as code, and executes it in a general-purpose interpreter to compute the answer. Each attempt leaves behind code and execution records that provide evidence for diagnosing failures, creating an opportunity for models to improve their solution strategies through reflection. To evaluate both computational visual reasoning and improvement through reflection, we introduce Code-with-Image Bench (CwI-Bench), a benchmark of 30 procedurally generated tasks with verifiable answers. Its disjoint training, validation, and test instances separate strategy refinement and selection from evaluation on unseen instances. We propose Reflection-Driven Refinement, a framework that develops and refines reusable solution strategies through reflection on prior attempts without updating model weights. Without reflection, Code-with-Image outperforms chain-of-thought reasoning across all nine backbones. Our framework further raises test accuracy from 41.3% to 70.1% on GPT-5.6-luna and from 32.6% to 64.6% on Qwen3.5-27B. The strategies also transfer across model scales and families.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.