ELSA: Evaluating localization of social activities in urban streets using open-vocabulary detection
Abstract
We introduce ELSA: Evaluating Localization of Social Activities, to our knowledge the first multi-label benchmark for detecting individual and group activities in still street-level images. ELSA draws on theoretical frameworks in urban sociology and design, and covers in-the-wild scenes in which group sizes and activity types vary significantly. The dataset includes 1,500 manually annotated images with more than 6,000 bounding boxes. Existing Open-Vocabulary Detection (OVD) models face a number of challenges. They often struggle with semantic consistency across diverse inputs and are sensitive to slight variations in input phrasing, leading to inconsistent performance. The calibration of their predictive confidence, especially in complex multi-label scenarios, remains suboptimal, frequently resulting in overconfident predictions that do not accurately reflect their context understanding. We expose these limitations and show that modern OVD models perform poorly on ELSA. We further propose Normalized Log-Sum-Exp (N-LSE), a context-aware confidence score for these models. Unlike the commonly used Max-Logit, which relies on a single token, N-LSE accounts for every label in the query, yielding a less inflated ranking and better detection results. Finally, we present Dynamic Box Aggregation (DBA), an evaluation algorithm that penalizes contradictory predictions for the same person or group and recovers correct boxes that NMS discards, providing a more reliable AP. We release our dataset and code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.