HQS-Tuning: A Hidden-State Quality-Guided Data Filtering Framework for Efficient Multimodal Instruction Fine-Tuning
Abstract
Instruction fine-tuning is a key technique for enhancing the performance of Multimodal Large Language Models (MLLMs), but massive and noisy datasets often incur prohibitive training costs with diminishing returns. Recent studies suggest that filtering a small amount of high-quality data for instruction fine-tuning can achieve highly efficient and superior training outcomes. However, existing multimodal data filtering approaches predominantly depend on external evaluation models or predefined proxy metrics, without leveraging the internal perception of the target Vision-Language Model (VLM) itself. This limitation often results in a mismatch between the filtering criteria and the actual learning requirements of the model being fine-tuned. To address these issues, we propose a novel perspective: the image-conditioned hidden states of VLMs implicitly reflect their evaluation of multimodal training data quality. Based on this insight, we propose an evaluator-guided two-stage data filtering framework. It extracts the target VLM's hidden states as representative features and builds a lightweight classification model using strict quality-tier supervision to score and select the optimal multimodal training subset. Our extensive experiments are conducted across two mainstream VLMs and eight diverse multimodal benchmarks. 172 In terms of data efficiency, our experiments demonstrate that fine-tuning on less than 1% of the large-scale LLaVA-665K dataset (merely 5K samples) enables our method to outperform models trained on the complete 665K dataset. 173 In comprehensive benchmark evaluations, our approach consistently surpasses multiple state-of-the-art multimodal data selection algorithms (such as COMPACT, ICONS, and Self-Filter) under the same data budgets, maintaining robust performance advantages across varied data scales.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.