VeriPixel: Code-as-Verification for Self-Improving Pixel Prediction
Abstract
Generative foundation models attract increasing attention for pixel prediction, as the image-to-image formulation offers the potential to unify image generation and pixel-level understanding tasks. However, existing prompt-based approaches typically rely on fixed, handcrafted prompts for single-round inference. They struggle with the limited semantic understanding capability of generative models, leaving a substantial performance gap relative to tuning-based methods. To this end, we propose VeriPixel, a verification-driven agentic framework for self-improving pixel prediction, which uses verification feedback to progressively enrich generation guidance with semantic information missing from the generative model. First, to provide reliable feedback for self-improvement, we propose code-as-verification, which uses a vision-language model (VLM) to generate executable code, transforming the vanilla single-pass VLM-as-a-judge into an iterative perception-execution-observation process. Instead of image-level VLM judgment, it provides iterative, pixel-grounded verification process. Nevertheless, producing reliable feedback through code execution requires an effective verification strategy. Therefore, instead of a single improvement loop, we propose bootstrapped self-improvement that first uses verifiable feedback from annotated training data to optimize verification strategies offline. With these strategies as instructions, the verifier provides reliable feedback and pixel evidence to bootstrap online self-improvement. Experiments show that VeriPixel improves upon direct prediction by 18.5 cIoU, 11.7 mIoU, and 3.4 dB PSNR across three representative pixel tasks, demonstrating significant gains from self-improvement and achieving competitive performance against tuning-based methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.