acceptodds
Under review as a conference paper at ICLR 2027

Layer-Specialized Model Merging for Training-Free Cross-Modal Reasoning Transfer

Abstract

Text-only reasoning LLMs are strong multi-step problem solvers, but transferring this behavior to multimodal LLMs (MLLMs) usually requires costly multimodal post-training and scarce verifiable reasoning supervision. We study a training-free alternative: fusing a text-only reasoning decoder with a homologous MLLM decoder while keeping the MLLM visual encoder and projector unchanged. This setting is more delicate than standard model merging because the fused decoder must integrate reasoning behavior from a text-only model while continuing to interpret visual inputs. We introduce (Attention-calibrated Norm-Orthogonal merging), a layer-specialized fusion method for cross-modal reasoning transfer. represents the two decoders as task vectors relative to a shared base model and computes closed-form, label-free fusion weights from task-vector salience. To account for depth-dependent visual-token use, it calibrates these weights with the MLLM's layer-wise visual-attention decay, using a continuous prior that favors the vision branch in shallow layers and the reasoning branch in deeper layers. Across homologous MLLM–reasoner pairs, improves reasoning-heavy multimodal benchmarks without gradient updates or labeled multimodal reasoning data. We also evaluate the boundary of this transfer: MME-RW perception results and MathVista performance reveal task-dependent tradeoffs between reasoning gains and visual capabilities. Together, the results suggest that layer-specialized merging is a practical route for training-free cross-modal reasoning transfer within the studied homologous model families.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.