acceptodds
Under review as a conference paper at ICLR 2027

Rewarding Non-Verifiable Intermediate Steps for Photorealistic Multimodal Reasoning

Abstract

Reinforcement learning for reasoning has shown strong results in mathematics and code generation, where intermediate steps can often be checked using deterministic rules or exact outcomes. In photorealistic multimodal reasoning, structured intermediate steps may also be needed, but they often lack a single exact answer and must instead be judged from visual and linguistic context. This makes it difficult to design reliable step-wise rewards for photorealistic multimodal reasoning. We propose a staged GRPO framework that first optimizes format and global accuracy rewards, then introduces step-wise rewards for intermediate reasoning. We consider two ways to specify the reasoning structure: one guided by human strategies and one proposed by a model from task examples. We introduce MOSAIC (Multimodal Object-Specific Ambiguity Inference from Context), a multilingual visual word inference testbed over photorealistic images that requires integrating visual and linguistic information and naturally elicits a structured reasoning process, which we decompose into four human-inspired steps. Experiments across 12 languages show improvements over prompt-based, supervised fine-tuning, and vanilla GRPO baselines, with the largest gains in lower-resource languages. Gains extend across Qwen and InternVL model families, and the larger models improve on three languages unseen during RL training. We further ask whether the reasoning structure must be manually specified. On SpatiaLQA, a distinct spatial-reasoning and action-planning task, we use model-proposed reasoning steps and evaluation criteria, yielding a 4.50-point improvement in F1-score over the baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.