BiVaR:Bidirectional Variational Discrimination for Cross-Model Reasoning Transfer
Abstract
Vision-language models (VLMs) inherit their language backbone from pretrained large language models (LLMs), yet still fall behind them on tasks requiring complex multi-step reasoning. Training-free reasoning transfer via model merging avoids costly multimodal supervision. However, existing merging strategies rely largely on parameter-level relations or individual response strength, which provide only indirect evidence of cross-model transferability and may therefore introduce model- or modality-specific interference during reasoning transfer. We introduce BiVaR, a distribution-guided framework that identifies transferable reasoning components through their functional consistency across models. Specifically, transferability is determined by whether a neuron's activation behavior remains compatible with the other model's activation distribution, while bidirectional evaluation retains only reasoning components with mutually preserved functional behavior across models. Reasoning transfer is then restricted to the corresponding shared-neuron subspaces, after adaptive low-rank decomposition preserves dominant cross-model directions while suppressing residual variations. Extensive experiments across multiple VLM-LLM pairs show that BiVaR consistently improves mathematical and text-only reasoning while preserving or improving general multimodal capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.