MIRROR: Learning from the Other View for Multi-Modal Reasoning
Abstract
Vision-language models (VLMs) can behave very differently depending on whether the same underlying problem is presented textually or visually. This inconsistency suggests that different views expose different reasoning paths and failure modes that standard multimodal post-training does not fully exploit. To study and exploit this phenomenon, we construct ODA-Data, a high-quality paired multimodal geometry dataset with text-dominant, image-dominant, and combined image+text views of the same problems, together with splits for training and evaluating view-dependent reasoning behaviors. We then developModality-Informed Reciprocal Reasoning Optimization (MIRROR), a reinforcement learning approach that derives a cross-view self-distillation signal from the model’s own modality-dependent capabilities. For each problem, MIRROR evaluates the model under all views, selects the best-performing view as a teacher, and trains the remaining views using a GRPO task-reward objective augmented with a reverse-KL regularizer toward that teacher. Across reasoning benchmarks that evaluate on multimodal mathematics problems, MIRROR improves over standard RL and yields more accurate and consistent behavior across different input modality representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.