Preserve the Language, Use the Vision: Training-Free Data Selection for Vision-Language Instruction Tuning
Abstract
Visual instruction tuning is essential for adapting vision-language models (VLMs) to vision-language tasks; however, this process may perturb the language representations acquired by their underlying large language models (LLMs) during language pretraining, thereby degrading their original language capabilities. Existing approaches primarily mitigate this issue through architectural modifications, additional text-only data, or training-time constraints, while largely overlooking the role of vision-language instruction data itself. In this work, we find that different vision-language samples induce substantially different degrees of perturbation to the original language representations of their underlying LLMs, suggesting that data selection can provide an effective way to mitigate language capability degradation. Based on this observation, we propose PLUV, a training-free vision-language data selection framework. PLUV first employs language forgetting risk estimation (LFRE) to select samples that induce less perturbation to the original language representations, and then uses visual dependency estimation (VDE) to measure the actual dependency of each sample on visual information. On text-only and vision-language benchmarks, PLUV achieves 98.6% and 102.0% of the baseline performance, respectively, demonstrating its ability to preserve language capabilities while maintaining strong vision-language performance and thereby achieving a better balance between the two.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.