acceptodds
Under review as a conference paper at ICLR 2027

Sperry: Multimodal Reasoning with Unified Model

Abstract

Current test-time inference in multimodal models largely relies on textual thinking traces, which are insufficient for reasoning processes that are hard to verbalize, such as spatial reasoning and visual geometry. Existing alternatives use external tools or latent-space reasoning to obtain visual thoughts, but they either lack flexible pixel-level generation or fail to expose interpretable intermediate reasoning. Unified multimodal models that generate both text and images offer a promising foundation, yet their post-training for rigorous multimodal reasoning remains underexplored due to the lack of high-quality interleaved text-visual reasoning data and effective training recipes. We introduce Sperry-43K, a dataset of interleaved text-visual reasoning traces across diverse visual reasoning tasks, together with SperryBench, a stratified evaluation benchmark. We further propose a complete post-training pipeline combining two-stage supervised fine-tuning with reinforcement learning. To reduce training-inference mismatch during RL, we introduce Generative Self-Prompted Policy Optimization (GSPPO). Our 14B model improves the base unified multimodal model from 19.1% to 51.9% on SperryBench, a held-out split of its own task generators, outperforming much larger public models such as Qwen3.5-397B and Gemma-4-31B evaluated zero-shot, as well as a latent-reasoning baseline trained on the same data, and providing a foundation for generative multimodal reasoning. Code and data will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.