Learning Joint Video and Text Representations and Taxonomy of Actions in Videos
Abstract
The video of a complex action can be decomposed into a temporally ordered sequence of simpler sub-actions, each of which can, in turn, be recursively decomposed to form a video tree. Given a video whose only annotation is a set of names of possible complex action contained in it, this paper aims to simultaneously segment the video into clips of the complex actions, learn a video tree of sub-actions for each segment, and assign linguistic names to its nodes. Each video tree has a co-existing text tree, formed by the text names of the video tree nodes. We refer to both tress together as the definition tree. In the absence of annotations, we generate text trees using LLMs. To construct video trees, we recursively divide estimated action segments in video into ordered temporal sub-actions. For both trees we compute representation features bottom-up, starting from the leaf actions. Action–sub-action attention (ASA) lets neighboring sub-actions exchange information before they are combined into a parent action in both trees. As a second objective, we use taxonomic supervision to draw semantically related actions close to each other to form a taxonomy of the complex actions in the input video. Compared with Ti-FAD, our video features better preserve the temporal order of sub-actions for unseen actions (+16.7%), yield better taxonomic grouping (+18.2%), and improve zero-shot localization mAP by 9.0-10.8% on THUMOS14 and 2.3-3.9% on ActivityNet.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.