acceptodds
Under review as a conference paper at ICLR 2027

VisualReasoningGym: Adaptive and Verifiable Datasets for Compositional Reasoning Across Scientific and Natural Images

Abstract

We experimentally identify a structural limitation of static visual datasets for training multimodal reasoning models for compositional reasoning with RLVR. Static datasets bind images and annotations to fixed questions whose complexity and learning value varies across model families and can decline during training. Yet the same images and annotations can support new questions that recover learning signal and increase the complexity of the visual reasoning task. We therefore argue that visual reasoning datasets should be more effectively represented as a generative process over the images, annotations, and reasoning operations, with questions generated at training time to match the policy, and at evaluation time at increasing difficulty until the model breaks. Building on these findings, we introduce VisualReasoningGym, a framework that repurposes existing static visual datasets and its metadata into compositional visual reasoning programs that factorizes a question into image dependent grounded atomic skills and a verifiable reasoning program. VisualReasoningGym samples multiple atomic skills per image as individual questions and composes their answers using parametric reasoning programs, ranging from aggregation and affine expressions to shortest-path and combinatorial-optimization problems. This decomposition provides exact final-answer verification and exposes intermediate outcomes. Rollout feedback adjusts compositional difficulty to the policy’s evolving capabilities. We evaluate VisualReasoningGym as (1) a source of on-policy RL training data and (2) as benchmark for measuring compositional proficiency: the mean number of consecutive difficulty levels passed by a model. Continuing RLVR from Qwen3-VL-8B raises held-out pass@5 compositional proficiency from 4.55 to 6.19 of 8 levels and accuracy by +22.3 points, approaching the compositional proficiency of frontier models evaluated zero-shot on the same in-distribution benchmark. Training also notably improves average performance on external visual reasoning benchmarks across Qwen3-VL and Nemotron-Omni, demonstrating notable positive transfer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.