GroundBricks: A Configurable Benchmark for Compositional Visual Grounding
Abstract
Visual grounding aims to localize image regions specified by language queries. Existing benchmarks, however, often focus on direct query–object matching, providing limited and systematic control over the diversity and complexity of grounding tasks. Constructing richer benchmarks typically requires additional region-level annotations, which are costly and difficult to scale. We argue that the required localization supervision already exists. Current object detection and instance segmentation datasets contain abundant verified region annotations. What is missing is a mechanism for composing these annotations into diverse and controllable grounding tasks. Based on this insight, we introduce GroundBricks, a configurable benchmark for compositional visual grounding, together with a fully automated construction pipeline. GroundBricks covers five capability categories: referential grounding, set grounding, attribute reading, relational chaining, and knowledge-grounded selection. It also spans three complexity levels. Our pipeline uses existing object detection annotations as its foundation, enriching annotated regions into structured semantic representations that serve as reusable grounding bricks for composing different tasks. The pipeline then applies rule-based query generation and validation. This process requires neither manually written queries nor additional region annotations. GroundBricks provides a fixed test set for reproducible comparison while supporting extensions through new composition rules and query templates. Its evaluation scope can therefore evolve alongside advances in model capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.