Rethinking Influence-Based Data Selection: A Curvature-Aware Sampling Framework for Visual Instruction Tuning
Abstract
Selecting informative training data is central to efficient multimodal instruction tuning. While existing methods widely rank examples via first-order, pointwise influence on validation loss, such filtering provides no guarantees on the validation risk after subset retraining. To address this limitation, we introduce a budget-constrained Poisson sampling framework that optimizes inclusion probabilities to minimize the expected validation risk after subset retraining. Crucially, we show that the validation curvature governs the retrained risk and derive a second-order risk surrogate endowed with a closed form optimal inclusion rule. We further establish an excess risk bound that decays as with the sampling budget , and characterize the relationship between existing influence function methods and the proposed framework. Extensive experiments across multiple vision language models, multimodal benchmarks, and selection budgets demonstrate the effectiveness of our method; notably, with an expected selection budget of only 6.4%, it retains 98.38% of the performance of training on full dataset.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.