ConfluxCount: Forming and Enumerating Semantic Confluences for Text-Guided Object Counting
Abstract
Text-guided object counting requires converting spatially distributed query evidence into one count per object, while retaining weak responses and avoiding duplicate counts. Existing density and detection readouts respectively preserve coverage or expose instances, but neither explicitly organizes distributed evidence into object-level counting units before enumeration. We introduce ConfluxCount, which instead counts semantic confluences: object-centered structures formed by a text-grounded semantic flow over a dense probe lattice. A crowding-adaptive determinantal point process (DPP) then enumerates these confluences, suppressing redundant probes while preserving nearby instances, and produces both counts and locations from a unified probabilistic readout. On CLOC-v1.1, the jointly trained model reduces MAE from 8.39 for Count Anything to 6.03. It also reaches 4.14 MAE on KubriCount with category text and 5.75 on MixCount with positive text in supporting evaluations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.