acceptodds
Under review as a conference paper at ICLR 2027

The MixCount Dataset: Exact Annotations at Scale for Open-Vocabulary Object Counting

Abstract

Object counting is a foundational vision task with over a decade of dedicated research, yet state-of-the-art models still fail systematically in the mixed-object setting that dominates real-world applications such as industrial inspection and product sorting. We show that this gap is strongly driven by limitations in existing training and evaluation data. Counting annotation is hardest exactly where it matters most, since a dense scene has to be labelled instance by instance, and existing datasets contain numerous errors. Existing synthetic alternatives remove the annotation cost but lack diversity and realism. We address this with MixCount, a dataset and benchmark for mixed-object counting whose annotations are exact by construction. An automatic pipeline synthesizes images, fine-grained textual descriptions and counting annotations at scale, eliminating the labeling ambiguity that plagues prior datasets. Evaluating state-of-the-art counting models on MixCount exposes severe degradation in the mixed-object setting. More importantly, training these models on our synthesized data yields substantial gains on real-world benchmarks, reducing MAE by 20.14% on FSC-147 and by 18.3% on PairTally. These results establish MixCount as both a benchmark and a training dataset for fine-grained counting, and demonstrate that exact annotations at scale help address a long-standing data bottleneck in counting models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.