Q-ViSTA: Semantic Identity and Query-Localized Transition Dynamics for Generalized Video Skill Discovery
Abstract
Generalized category discovery (GCD) enables open-world systems to organize known and novel categories using partial supervision. Extending this capability from images to videos requires preserving not only semantic identity, but also the localized interactions that define an operation. The same object can participate in different skills, yet global video aggregation can dilute the regions where their defining changes occur. To address this problem, we propose Q-ViSTA, which combines Semantic Identity with Query-Localized Transition Dynamics. Following a where-before-how design, recurrent Action Queries first collect regional trajectories, and a multiscale transition operator then describes their evolution. A parallel semantic branch retains global context, while nonlinear fusion brings both representations into a shared discovery space after temporal modeling. This representation supports both non-parametric clustering and a jointly learned parametric classifier without requiring region annotations. Experiments across six human-action and robot-skill protocols demonstrate improved domain-level discovery over SimGCD for both variants under a shared frozen backbone, with pronounced gains on unlabeled multi-step tasks that combine familiar atomic skills.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.