What to Count Before How Many: Text-Responsive Zero-Shot Object Counting via Two-Stage Regression
Abstract
Integrating text guidance into class-agnostic object counting has drawn growing interest. Historically, lightweight regression-based paradigms (e.g., CLIP-Count) have dominated the field due to their high inference efficiency and accuracy on dense, overlapped objects. However, standard regression architectures suffer from a fundamental, overlooked flaw: they tend to fall into visual pattern recognition, marginalizing text constraints entirely. Due to conventional evaluation metrics (e.g., Mean Absolute Error) assess only counting accuracy, this severe text-ignoring issue has long been obscured by over-counting on visual textures. To serve the need of overall counting accuracy, recent mainstream research has increasingly shifted toward heavy detection-based frameworks, which happens to be text-aware by its original purpose. Nevertheless, such detection frameworks enforce explicit target localization at the cost of high computational overhead. To bridge this gap, we revisit the regression-based paradigm and propose a two-stage framework that restores robust text-awareness without compromising efficiency. Stage 1 extracts fine-grained cross-modal features to generate a prompt-compliant mask corresponding to the specified objects. Stage 2 utilizes this mask as a spatial prior to generate a similarity map for density map regression, the density map will be used for the final counting process. Extensive experiments demonstrate that our approach eliminates text-ignoring issue while maintaining high computational efficiency and competitive counting accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.