Beyond Actions and Objects: Video Classification in the Wild
Abstract
Existing video classification datasets are largely built around predefined actions, objects, or their combinations. However, video classification in the wild is more challenging: videos within the same category can vary widely in subjects, scenes, and appearance, while different categories can be visually similar yet semanti-cally distinct. To study this problem, we introduce EventTree, a human-in-the-loop framework that discovers and organizes event categories from large-scale video collections through top-down taxonomy construction. Applying EventTree yields WildEvent-1K, a hierarchical event-centric video classification dataset with 538K videos and 1, 040 fine-grained event categories in a three-level taxonomy. WildEvent-1K exhibits substantial intra-class variation, subtle semantic differences between neighboring categories, and a long-tailed distribution, providing a challenging testbed for real-world classification. These characteristics make conventional video encoders with direct classification heads less effective, motivating classification through next-token prediction with a VLM pretrained across diverse visual domains. To accommodate the hierarchical structure of WildEvent-1K, we propose Generative Hierarchical Video Classification (GHVC), which predicts categories level by level, restricting each prediction to valid children of its predicted parent. Experiments reveal a clear contrast with conventional datasets such as Kinetics-400: specialized video encoders remain strong in action recognition but are substantially less effective on WildEvent-1K, where VLM-based classification performs better. GHVC further improves it through hierarchical next-token prediction, demonstrating its effectiveness for video classification in the wild. Models, code, and data will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.