OnceSelect: A Reusable Data Selection Framework for Efficient Multimodal Instruction Tuning
Abstract
Multimodal instruction tuning is the de facto training method for adapting multimodal large language models (MLLMs), yet the large-scale image-text datasets it relies on are often highly redundant. Effective data selection is therefore essential for efficient multimodal instruction tuning. However, existing selection methods are often tightly coupled to the specific dataset, target model, or training state used to derive their selection signals. Consequently, applying them to a new dataset or downstream model typically requires retraining the selector or recomputing model-dependent scores, leading to substantial repeated computation. This raises a fundamental practical question: can a selection signal be learned once and reused across datasets and target models? To this end, we propose **OnceSelect**, a reusable data selector for multimodal instruction tuning that is trained only once and then directly applied to unseen datasets. OnceSelect encodes image-instruction pairs in a frozen joint multimodal space, clusters them to obtain pseudo-labels that capture the coarse semantic structure, and trains a lightweight selector and retains an intermediate checkpoint selected according to a fixed validation macro-accuracy target. Samples receiving low confidence under this learned semantic partition are treated as potentially informative candidates. Because the selector is independent of the downstream MLLM, it can be reused across candidate datasets without retraining or model-dependent scoring, while each selected subset can subsequently be reused across model architectures. Using only 15% of LLaVA-625K, OnceSelect retains 99.2% of full-data aggregate performance across nine evaluation metrics. Without selector retraining, it achieves 103.4% and 102.5% relative performance on the unseen Vision-Flan-186K and LRV-Sub-180K, respectively. Moreover, the same 93.7K LLaVA subset achieves 99.9–106.7% relative performance across the evaluated Qwen3-VL and InternVL3 models without model-specific scoring or reselection. Together, these results establish OnceSelect as a reusable and efficient data selection framework across datasets and downstream model architectures. Code and the trained selector will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.