Counting the Uncounted: Correcting Occlusion-caused Bias in Crowd Counting Benchmarks
Abstract
Vision-based crowd counting has advanced considerably, yet its real-world reliability remains limited by occlusion, which existing methods treat as undifferentiated noise rather than as a structured scene element. We challenge a foundational assumption of current benchmarks: that ground-truth (GT) annotations reflect true crowd size. By auditing four widely used datasets (ShanghaiTech A, UCF-QNRF, JHU-Crowd++, and NWPU-Crowd), we introduce a three-class occlusion taxonomy: Atmospheric (AO), Barrier (BO), and Embedded Occlusions (EO). We identify EOs as the most consequential yet overlooked category. EOs are small, non-transparent objects embedded within the crowd, such as umbrellas, picket signs, and flags. Occlusions appear in 40.73% of images, and 78.46% of these involve EOs, yet no annotation protocol accounts for the individuals concealed behind them. We correct the GT through occluder-specific annotation and local density inference with a new annotation tool, Occlusion Density Fusion (ODFlow). The correction shows that existing GT undercounts EO-affected images by 6.3% on average, so published MAE understates real-world error by about 22 counts per image. Beyond data correction, we propose ODFNet, an occlusion-aware counting model. Its Density Reconstruction and Count Fusion (DRCF) module reconstructs human and EO density maps separately, then fuses them while removing redundancy introduced by overlapping EOs. Validation on both synthetic crowd scenarios and real crowd images shows that ODFNet recovers the ground-truth count regardless of EO type, number, or spatial pattern, whereas existing methods drift substantially. We release corrected GT for all four datasets, a comprehensive benchmark re-evaluation, a newly curated occlusion-aware dataset (ODF-A), and ODFNet.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.