PuzzleCodeBench: A Visual-to-Code Benchmark for Executable Puzzle Solving
Abstract
Visual puzzles offer a challenging testbed for vision-language models (VLMs), requiring visual parsing, rule grounding, structured state tracking, and verifiable solution search. However, existing benchmarks cover limited scales and complexities, and typically evaluate models via natural-language answers. Such non-executable, instance-specific responses make it difficult to verify whether models learn reusable solving strategies, limiting the evaluation of scalable generalization and efficiency. To address these limitations, we introduce **PuzzleCodeBench**, which reframes puzzle solving under a visual-to-code paradigm: given a puzzle image and rules, models must parse the visual board and generate a complete Python solver that computes the solution. The benchmark comprises 790 instances across 30 grid-based logic puzzle families and grids up to 2525. Based on this collection, we develop an execution-centered pipeline with a standardized observation interface to evaluate models along four axes: visual perception, solver correctness, cross-instance generalization, and algorithmic efficiency. Experiments across frontier VLMs show that executable puzzle solving remains challenging, with model-specific bottlenecks in visual parsing and solver synthesis. Meanwhile, even strong models show substantial performance drops on larger and harder held-out instances. We further apply execution-grounded reinforcement learning to Qwen3.5-9B on 11 puzzle families, achieving a 2.5 improvement in execution accuracy and zero-shot generalization to 9 unseen families. These results highlight **PuzzleCodeBench** as both a diagnostic benchmark and a training environment for robust code-based puzzle solving.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.