Balancing Cross-Modal Class Alignment and Local Separation for Zero-Shot Skeleton Action Recognition
Abstract
Zero-shot skeleton action recognition requires visual–semantic relations learned from seen classes to generalize to unseen actions. Existing cross-modal methods mainly reduce instance- or distribution-level modality discrepancy, but such alignment does not directly guarantee a suitable class structure. Visual and semantic representations of the same action may still form inconsistent class centers, while stronger positive alignment may compress the local separation between similar actions. We address this alignment–separation problem with a hierarchical class-geometry objective. Cross-Modal Prototype Neighborhood Modeling (CPNM) maintains visual and semantic prototype memories and uses modality-swapped anchors to form positive cross-modal class neighborhoods. Hard-Neighbor Discrepancy Amplification (HNDA) then identifies the nearest confusing class prototypes and imposes a relative margin around these neighborhoods. Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD demonstrate competitive performance under ZSL and GZSL protocols. Controlled ablations and geometry analyses further verify the complementary roles of cross-modal neighborhood alignment and local hard-neighbor separation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.