acceptodds
Under review as a conference paper at ICLR 2027

DataValve: Selecting-while-Training for Efficient Visual Instruction Tuning

Abstract

Visual instruction tuning has become increasingly expensive as multimodal large language models (MLLMs) and training corpora scale. Existing data selection methods mostly follow an offline select-then-train pipeline: they score a large candidate pool, often with an additional target MLLM traversal, and then fine-tune on a fixed subset. This pipeline is both costly and blind to the evolving training state, so it cannot avoid samples whose marginal utility has already decayed. We propose DATAVALVE, a selecting-while-training framework for fixed-pool visual instruction tuning. A lightweight router selects training samples from pre-extracted features and adapts through delayed golden-set loss feedback. To counter the tendency of loss-only feedback to repeatedly favor high-loss sources, an online LoRA-influence gate modulates credit using relative update strength. This coupling prioritizes currently useful samples and suppresses redundant low-impact updates without external teachers, downstream validation sets, or explicit diversity losses. Experiments on LLaVA-v1.5-7B show that DATAVALVE routes only 20% of training data through target MLLM training while achieving 98.4% of full-data fine-tuning performance, outperforming strong baselines under the same routing budget with the best efficiency–accuracy balance. Further experiments demonstrate transfer across candidate pools, model scales, and MLLM families. Our code is included in the supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.