acceptodds
Under review as a conference paper at ICLR 2027

Prior-Guided Importance Weighting for Supervised Fine-Tuning of Multimodal Large Language Models

Abstract

Supervised fine-tuning (SFT) is critical for adapting pretrained multimodal large language models (MLLMs) to downstream tasks, yet it is often performed with limited target-domain data, making the training outcome highly sensitive to how supervision is utilized. Recent data-centric fine-tuning methods address this challenge by selecting or reweighting target samples, but they often rely on restrictive assumptions or additional components, and their emphasis on better fitting the target distribution may still lead to over-reliance on scarce target supervision. In this work, we view limited-data MLLM fine-tuning as a trade-off between target adaptation and pretrained prior preservation, and propose Prior-Guided Importance Weighting (PGIW), an importance-sampling-based framework that incorporates pretrained model guidance without accessing the original pretraining data. Our theoretical analysis derives an optimal balancing coefficient under an expected Kullback-Leibler divergence risk, while importance weighting enables PGIW to approximate the unavailable pretraining-distribution term using target-domain samples. We further instantiate PGIW as an efficient fine-tuning algorithm for MLLMs. Experiments on multimodal benchmarks demonstrate that PGIW improves target-task performance and cross-domain generalization over standard fine-tuning baselines. The source code is available at https://anonymous.4open.science/r/2026PGIW-67E3/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.