GapSeer: Hyperbolic Semantic Axis Responses for Weakly Supervised Fine-Grained Video Anomaly Detection
Abstract
Existing weakly supervised fine-grained video anomaly detection (WS-FVAD) methods typically construct category responses based on cross-modal semantic similarity between concrete event instances and category texts describing abstract anomaly semantics. However, the semantic granularity mismatch between abstract category descriptions and diverse visual realizations makes similarity-based responses unreliable. Although hyperbolic entailment cones provide an asymmetric mechanism for modeling instance-to-category relations, visually similar patterns shared by different anomaly categories may cause an event instance to fall into multiple category cones, resulting in ambiguous category responses.To address this issue, we propose GapSeer, which constructs fine-grained category responses by modeling how strongly a video segment specializes along each category-specific semantic direction, rather than relying on overall visual-text similarity or cone inclusion alone. Specifically, the Hyperbolic Semantic Axis (HSA) defines a directed semantic trajectory from abstract category semantics toward concrete event realizations, while the Signed Axial Response (SAR) quantifies the signed progression of each video segment along the corresponding axis. This formulation accommodates diverse visual manifestations within the same anomaly category while distinguishing visually similar but semantically different events, thereby reducing ambiguity among competing category responses. Extensive experiments on the XD-Violence and UCF-Crime benchmarks demonstrate the effectiveness of our method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.