acceptodds
Under review as a conference paper at ICLR 2027

FactorGraph-RLVR: Diagnosing Reward Informativeness in Factorized Visual Reasoning Curricula

Abstract

Reinforcement learning with verifiable rewards (RLVR) has improved multimodal reasoning, but gains attributed to curriculum design may instead reflect how often training produces informative group-relative updates. We introduce \ours, a controlled visual-graph testbed that crosses visual masking and reasoning load in a matched factorial design, together with a parser-gated dense verifier. To separate curriculum composition from optimization opportunity, we compare reasoning-heavy and uniform curricula under a fixed optimizer-step budget, match checkpoints by cumulative informative updates, and ablate dense supervision with exact-only rewards. A preregistered evaluation covers 28 frozen model checkpoints, three final evaluation splits, and four decoding conditions. Under greedy decoding, reasoning-heavy dense RLVR improves exact success over the unadapted model on all three splits. The gains are 4.1, 4.2, and 4.7 percentage points on in-distribution, style-shift, and length-shift data, respectively. All 36 seed–split–decoding contrasts are positive, and every 95% bootstrap confidence interval computed from paired source groups excludes zero. Uniform training yields smaller gains, informative-update matching retains most of the reasoning-heavy effect, and exact-only training remains near the base model. These results provide controlled evidence that curriculum composition influences visual RLVR beyond optimizer-step budget and measured informative-update opportunity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.