acceptodds
Under review as a conference paper at ICLR 2027

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

Abstract

Vision-language models (VLMs) advance rapidly, but fixed benchmarks can saturate and constructing targeted stress tests remains labor-intensive. Collecting and annotating existing images does not ensure that they satisfy task-specific visual conditions or challenge current models. To solve this, we present SABRE, a scalable, semi-automated pipeline that converts Test Primers (Markdown Task Designs with Data Schemas) into structured specifications, generated or edited images, and question–answer pairs. A Filtering VLM screens candidates for obvious generation defects and prunes readily solved cases, focusing human review on challenging candidates. Reviewers verify task validity, correct annotations, and perform localized image repair when needed. As our main case study, we instantiate SABRE-Prior to test whether VLMs follow visual evidence over world priors—learned expectations about familiar objects and scenes. Its 600 images and 1,000 questions span four subsets: Context, Texture, Attribute, and Language Elicitation. Across six VLMs, macro-average accuracy ranges from 17.8% to 31.3% (22.6% mean). Pilot instantiations on Counting and Spatial reasoning, aided by procedural scene extensions, further demonstrate how the shared workflow generalizes to settings requiring precise quantitative and geometric control. Together, these components provide a reusable framework for constructing and refreshing targeted VLM stress tests as model capabilities evolve.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.