Let Clues Unfold: A Fine-Grained Multi-Label Dataset, Benchmark, and Dynamic Clue Mining Method for Explainable Harmful Short Video Detection
Abstract
Short videos have become a major channel where harmful content reaches the public, and a single video often mixes several harm categories through multimodal cues. Existing harmful video detection techniques fall short in two ways: harmful-content taxonomies remain inconsistent across platforms and studies, and public datasets are small and coarsely labeled; multimodal fusion methods do not resolve the semantic gap and temporal misalignment across modalities, and fail to provide frame-specific causal reasoning for model predictions. To overcome these problems, we construct FineHM, a fine-grained multi-label harmful short video dataset, spanning 8 categories with harmful frames and bounding boxes, to provide frame- and region-level evidence for explainable multi-label classification. Built upon this dataset, we establish a standardized benchmark including fixed data splits, evaluation metrics, and comprehensive baseline results to facilitate future research. Moreover, we propose ExplainHarm, a dynamic clue mining explainable harmful short video detection method, to surface dynamic evidence from category features and video content via a prompt-chained clue-mining process, and to achieve cross-modal semantic and temporal alignment through dual-scale temporal modeling. Experiments and statistical analyses demonstrate the dataset FineHM and the model ExplainHarm's superiority, achieving an 11.3% accuracy gain over SOTA multi-label models and a 0.6% IoU improvement over SOTA open-vocabulary object detection models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.