acceptodds
Under review as a conference paper at ICLR 2027

Dynamic Modality Scaling for Vision-Language Model Alignment

Abstract

Recent advances in vision-language models have achieved remarkable results in making language models understand visual inputs. However, a unified approach to aligning these models remains a challenge. In LLM-centric VLMs, language dominance is particularly plausible, because the decoder holds a strong prior over answers and explanations. Yet visual dominance is also possible in OCR or scene-heavy settings, where noisy patches overwhelm a simple question. The correct intervention must therefore be conditional, rather than globally image-boosting or text-suppressing. In this paper we propose dynamic modality scaling (VL-DyMS), a framework built on two frozen unimodal teachers and a multimodal student. Each teacher is obtained by masking one input: the text teacher sees only the question, the vision teacher only the image. The student is supervised by an answer loss and by a modality-aware distillation loss against each teacher. A dynamic controller adjusts the contribution of textual and visual knowledge at every step, so supervision follows the modality the student is currently underusing rather than a fixed schedule. We evaluate VL-DyMS on five datasets using Qwen2.5-VL and LLaVA-1.5. It shows an improved performance over the pre-trained baseline and standard LoRA adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.