acceptodds
Under review as a conference paper at ICLR 2027

Beyond Semantic Alignment: Towards Dense Instance Perception in Vision-Language Models for Open-World Counting

Abstract

Vision-Language Models (VLMs) achieve strong semantic understanding through large-scale image-text alignment, yet their ability to distinguish and enumerate nearby visual instances degrades substantially as object density increases. We investigate this semantic-to-instance perception gap through a unified language-guided counting benchmark built from four existing counting datasets and a controlled evaluation of representative visual encoders. Our analysis reveals that stronger semantic representations benefit counting mainly in sparse scenes, while their advantages diminish under increasing crowding. To alleviate this limitation, we propose DIP-CountVLM, a Dense Instance Perception Counting VLM for open-world scenarios. DIP-CountVLM combines Instance-Aware Feature Fusion (IAFF), which adaptively aggregates hierarchical visual features to recover instance-sensitive cues, with Dense Instance Token Reconstruction (DITR), which leverages image guidance to reconstruct spatially denser visual tokens. Extensive experiments show that DIP-CountVLM improves overall counting performance over a SigLIP2-based VLM baseline, with the most consistent gains in low- and medium-density regimes. Further analyses demonstrate that hierarchical feature fusion is most effective when coupled with dense token reconstruction, while extremely dense scenes remain challenging. These findings provide a systematic analysis of the visual representation bottlenecks underlying language-guided counting and a practical approach for alleviating them.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.