PixelProof: Visual Question Generation through Forward–Inverse Agreement
Abstract
From pre-training on human-written tokens to reinforcement learning on verifiable rewards (RLVR), obtaining cheaply verifiable data is the fuel for current large language models. RLVR often works in mathematics and similar domains because the labels are verifiable by construction and establish a strong self-supervision loop that bootstraps model capabilities. Constructing this loop outside math domains, especially in vision, is an interesting open challenge. We introduce PixelProof, a method that takes human-input prompts and automatically synthesizes novel visual-question images via Python where labels are verifiable by the generation loop itself (without human supervision). That is, PixelProof parameterizes a set of questions by 3 Python programs: (1) a Forward program that translates scene specifications into an expected answer; (2) a Renderer that renders an image; and (3) an Inverse program that progressively transforms the input image using image processing tools to derive its own verification answer independently. Comparing the answers between the Forward (what is designed) vs. Inverse programs (what is observed) self-supervises the generation. Generation can be steered toward (a) a human-defined topic e.g., music; (b) where in the image the answer's evidence lies; and (c) harder questions. As an initial demonstration of PixelProof, five coding agents working across nine visual reasoning directions produced over 2,400 verified questions. Fine-tuning three open-weight models on the newly generated questions improves their performance on held-out questions by an average of +10.7 points, and by +1.9 points on average across 16 external benchmarks. PixelProof shows a promising loop for automatically synthesizing visual questions for evaluating and training vision-language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.