Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
Abstract
Multimodal instruction tuning can be compute-inefficient when a fixed training budget is spread across large image–video pools whose examples vary substantially in utility. We present Goal-Driven Data Optimization (GDO), a framework that computes six sample descriptors for each candidate and constructs optimized 1x training subsets for different goals. Under a fixed one-epoch Qwen3-VL-8B-Instruct training and evaluation protocol on 32 H20 GPUs, GDO reaches the benchmark-specific Uni-10x references with fewer training examples and attains higher endpoint accuracy. Each reference is the lowest-scoring endpoint among four historical 512k-sample uniform branches, making the comparison descriptive rather than a matched causal estimate. GDO reaches these references after 35.4k samples on MVBench, 26.6k on VideoMME, 27.3k on MLVU, and 34.7k on LVBench, while improving endpoint accuracy by 1.38, 1.67, 3.08, and 0.84 percentage points, respectively. The gains are largest on MVBench and MLVU and smaller on LVBench, consistent with weaker alignment between its ultra-long-video setting and the short-video and image-dominant training pool. The four profiles span low-loss, coverage-oriented, and temporally focused allocations; among them, Temp+ gives the strongest overall endpoint profile.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.