CLIVO: Differentiable Cluster-Based Instruction Mixing with Learned Validation Objectives
Abstract
Multimodal instruction tuning depends on both the quality of individual examples and their composition in the training set. Target-specific data selection must therefore decide which examples are useful and how much of each type of supervision to include. Sample-level rankings leave this composition implicit, while group-level optimization faces three connected challenges: fine-grained units enlarge the search space, predefined validation sets may obscure target-relevant feedback, and fitted mixture-response predictors introduce additional approximation error. We propose CLIVO, which connects semantic allocation, automatically learned validation objectives, and differentiable mixture search. Semantic clusters give distinct instruction contents independently adjustable training shares. To evaluate these choices, CLIVO relates individual examples' merged-proxy losses to real-training probe scores, then selects and weights validation examples to form a target-specific objective. Differentiating this objective through reusable expert-merged proxies updates all shares jointly without a separate response predictor. Selecting 100K instructions from an 8M pool, CLIVO improves the four-target average over the best-performing state-of-the-art data selection baseline by 1.55 points; combining its quotas with ICONS increases this advantage to 1.94 points. Ablations show that moderate-budget experts preserve useful feedback, finer allocation benefits most targets, and sample-level weighting exploits target-relevant response differences obscured by cluster averaging. Semantic visualizations and reusable, modular components further support interpretable and flexible target-specific data selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.