The Value of Training Data Is Not Static: Optimization-Aware Adaptive Selection for Multimodal Instruction Tuning
Abstract
Multimodal instruction tuning relies on large‑scale image‑text datasets, making effective data selection critical for improving training efficiency. Existing approaches mostly adopt static offline scoring or measure visual dependence at inference time, yet these metrics cannot directly quantify how a training example’s gradient update will benefit the current evolving multimodal model. To address this gap, we propose **DOVAS**, a region‑based dynamic data‑allocation framework that combines offline visual‑evidence scoring with online optimizer‑aligned utility estimation. DOVAS first partitions the instruction corpus into fixed, semantically stratified regions. During training, it periodically re‑weights regional sampling probabilities by estimating the predicted impact of candidate optimizer updates without modifying region membership. This design lets the model draw adaptively useful training data throughout the optimization trajectory. Experiments on standard multimodal instruction‑tuning benchmarks demonstrate that DOVAS consistently outperforms competitive static data‑selection baselines. Using only half of the full‑data training compute budget, it surpasses the performance of full‑dataset training across multiple model families. These results demonstrate that the training‑time value of multimodal samples is not a static intrinsic property but depends on model state and optimizer dynamics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.