BindSplit: Benchmarking Seen and Unseen Compositional Generalization in VLMs
Abstract
Compositional generalization requires vision-language models (VLMs) to understand novel combinations of familiar concepts, and it is measured with compositional benchmarks. However, if most evaluation samples contain combinations that models already saw during training, high benchmark scores may not reflect generalization to new compositions. In this work, we ask whether compositional benchmark scores reflect generalization to unseen compositions. We focus on attribute–object bindings and introduce the BindSplit protocol, which splits the samples of existing benchmarks into seen and unseen by whether their bindings appear in a reference corpus that approximates the training data. We find that existing benchmarks contain very few samples with unseen bindings: at most 88 per benchmark against CC12M. To address this gap, we introduce BindSplit Bench, a benchmark built to compare performance on seen and unseen bindings directly. Its seen and unseen splits have similar sizes, share their attributes, and are generated and human-reviewed through the same pipeline. Across 75 VLMs, accuracy drops from seen to unseen bindings for 72 models, and seen accuracy, although correlated with unseen accuracy, overstates how much models improve on unseen bindings. Our results suggest that progress on existing compositional benchmarks may not translate directly to unseen compositions. By separating seen and unseen bindings, BindSplit Bench makes this generalization gap measurable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.