Layer-Specialized Model Merging for Training-Free Cross-Modal Reasoning Transfer
Abstract
Text-only reasoning LLMs are strong multi-step problem solvers, but transferring this behavior to multimodal LLMs (MLLMs) usually requires costly multimodal post-training and scarce verifiable reasoning supervision. We study a training-free alternative: fusing a text-only reasoning decoder with a homologous MLLM decoder while keeping the MLLM visual encoder and projector unchanged. This setting is more delicate than standard model merging because the fused decoder must integrate reasoning behavior from a text-only model while continuing to interpret visual inputs. We introduce (Attention-calibrated Norm-Orthogonal merging), a layer-specialized fusion method for cross-modal reasoning transfer. represents the two decoders as task vectors relative to a shared base model and computes closed-form, label-free fusion weights from task-vector salience. To account for depth-dependent visual-token use, it calibrates these weights with the MLLM's layer-wise visual-attention decay, using a continuous prior that favors the vision branch in shallow layers and the reasoning branch in deeper layers. Across homologous MLLM–reasoner pairs, improves reasoning-heavy multimodal benchmarks without gradient updates or labeled multimodal reasoning data. We also evaluate the boundary of this transfer: MME-RW perception results and MathVista performance reveal task-dependent tradeoffs between reasoning gains and visual capabilities. Together, the results suggest that layer-specialized merging is a practical route for training-free cross-modal reasoning transfer within the studied homologous model families.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.