VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
Abstract
Native visual reasoning treats visual states (i.e., images and videos) not as merely inputs to be understood or outputs to be rendered, but the intermediate steps through which a problem is solved. We introduce **VBVR-Pro**, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. **1) Task scaling.** VBVR-Pro curates a task space of *300* procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external video reasoning benchmarks. Analysis suggests that these gains reflect visual reasoning rather than instruction-pattern fitting. **2) Verifiable rewards.** VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. While the prevalent VLM-as-a-judge paradigm shows recurring failures, deterministic scorers align more closely with human judgments, and can further serve as reliable reward signals for large-scale multi-task reinforcement learning. **3) Modality study.** VBVR-Pro enables comparisons across leading image, video, and interleaved generators under a shared task distribution and scoring protocol. Our analysis shows that video generation remains strongest for modeling spatiotemporal persistency, while interleaved generation provides a compute-efficient alternative where visual states are crucial. We release all data, models, scorers, and code to facilitate future research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.