acceptodds
Under review as a conference paper at ICLR 2027

CounterCount: Diagnosing Knowledge Bias in How Vision-Language Models Count Attributes

Abstract

Do Vision-Language Models (VLMs) count what they see, or what they know about the world? We introduce CounterCount, a diagnostic dataset of paired factual and counterfactual (CF) images, in which the CF image contains an atypical count of a subject's attribute (e.g., a rabbit with four ears), together with annotations localizing the counted attributes. Humans answer both types of images nearly perfectly (98.9% and 97.5%), but across 16 VLMs (14 open-weight and 2 closed-source), average accuracy drops from 91.6% on factual to 49.5% on CF images, with most errors stemming from responses that match the canonical count, therefore revealing a bias towards knowledge priors rather than reliance on visual information. Based on these findings, we perform an analysis that reveals two interesting insights behind this gap. First, models exhibit a "System-1"-like response behavior, falling back on canonical counts after a shallow recognition of the subject. Second, model activations partially contain the visual evidence necessary to answer correctly, yet models underutilize it, attending less to the counted attributes when they provide biased answers. Input manipulations that break visual context, introduce cueing or reasoning, or direct amplification of the model's internal attention to relevant regions can reduce such biases. CounterCount goes beyond showing that VLMs fall back on knowledge priors, revealing that once a familiar context is recognized, models favor prior knowledge over available visual evidence rather than failing to encode that evidence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.