acceptodds
Under review as a conference paper at ICLR 2027

BIG: Bidirectional Information Gain for Visual Instruction Data Selection

Abstract

Existing visual instruction data selectors assess quality, coverage, or input contributions without jointly evaluating how the image and the particular instruction support the same reference answer. They can therefore retain examples whose answers require the image but are weakly constrained by the instruction, or follow from the instruction without requiring visual evidence. We propose BIG (Bidirectional Information Gain), which measures both conditional contributions using one frozen vision-language scorer. Each gain is the relative reduction in reference-answer loss from adding one input given the other. To align gain scales across sources and directions, BIG converts them to within-source percentiles and ranks examples by their minimum under fixed source quotas. This makes the weaker contribution decisive, preventing a strong direction from compensating for a weak one. BIG consistently outperforms random and single-direction selection in aggregate performance across two model–data settings, and also leads the comparison with established selectors under a shared training recipe. On Vision-Flan with Qwen2-VL-2B, BIG exceeds full-data aggregate performance across seven benchmarks using 50% of the data. Ablation analysis shows that removing calibration or allowing compensation shifts selection toward examples with weaker support in one direction. Selected-data analysis shows that BIG systematically recovers examples with stronger contributions in the direction overlooked by single-direction criteria.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.