acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Modality Interference for Unified Reasoning and Perception in Multimodal Large Language Models

Abstract

Recent text-only reasoning models show that System-2 reasoning can greatly improve math and code, but multimodal large language models (MLLMs) still rely on fast visual intuition with limited visuospatial gains. A dominant textual-proxy transfer recipe naively mixes supervision by converting images into caption-like intermediates and jointly training distilled long-form reasoning traces with multimodal perception supervision under a single objective. To address this gap, we conduct a systematic empirical analysis and attribute the trade-off to modality interference, where reasoning and perception objectives produce conflicting gradients under naive joint optimization. Building on this diagnosis, we propose UPRSD (Unified Perceptual and Reasoning via Staged Distillation), a training recipe that mitigates modality interference through decoupling, unification, and verifiable-reward reinforcement optimization, thereby improving reasoning without sacrificing perceptual grounding. Across multiple textual and multimodal benchmarks and different MLLMs scales, UPRSD strengthens multimodal and text reasoning while maintaining strong general multimodal understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.