RankRepair: Decoupling Record Priority and Budget Allocation for Training-Free Multimodal Data Selection
Abstract
Training vision-language models (VLMs) often requires substantial computational resources and large-scale data, making the selection of compact, high-quality instruction datasets essential for efficient multimodal learning. Selecting such subsets requires balancing individual record priority against the effective allocation of a limited data budget across the overall candidate pool. However, incorporating diversity or global structure into record-level scores does not by itself ensure adequate regional coverage in the selected subset. Regions can remain underrepresented after selection, motivating an explicit adjustment of subset-level allocation. To address this challenge, we introduce RankRepair, a novel two-stage framework that decouples record priority estimation from subset-level allocation adjustment. In the first stage (Rank), complete records are scored by integrating image–instruction alignment with local visual diversity to select an initial high-priority subset. In the second stage (Repair), the framework measures regional shortfalls relative to pool-proportional coverage floors and systematically mitigates them through record swaps while keeping priority evaluations fixed. By imposing explicit bounds on both the number of allowed swaps and the maximum priority loss per swap, RankRepair limits departure from the initial quality-based selection, ensuring each accepted swap strictly reduces the aggregate coverage deficit by one unit. Extensive evaluations on LLaVA-1.5 Mix-665K demonstrate that a 100K training subset selected by RankRepair outperforms the evaluated equal-budget baselines in mean performance across ten multimodal benchmarks. The selected subset retains 96.76% of the full-data mean performance, exceeding the strongest equal-budget baseline by 3.16 percentage points, supporting bounded coverage adjustment as an effective complement to quality-based ranking for visual instruction data selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.