acceptodds
Under review as a conference paper at ICLR 2027

Zero-Shot Skeleton-Based Temporal Action Segmentation

Abstract

Human action segmentation in continuous skeleton streams relies heavily on closed-set annotations, struggling to generalize to novel categories in the wild. While zero-shot learning (ZSL) offers a promising paradigm, existing skeleton-based ZSL methods are strictly confined to clip-level recognition. By encoding entire actions into global semantic prototypes, they inherently erase the fine-grained, localized motion primitives shared between seen and unseen categories, rendering them ineffective for dense frame-level segmentation. In this paper, we formalize and tackle the novel task of zero-shot skeleton-based temporal action segmentation. We propose SkelAtom, a framework that achieves compositional generalization by explicitly decomposing untrimmed actions into sparse combinations of shared motion atoms. To ensure these atoms capture transferable physical movements, we introduce Kinematic Anchor Grounding (KAG), which supervises the latent space with class-agnostic kinematic descriptors derived from body-part dynamics. Furthermore, a Cross-Modal Atom Codebook aligns both visual frames and textual prototypes into the same atomic basis, enabling visual knowledge learned from seen actions to transfer to unseen categories. We establish four zero-shot benchmarks on the MCFS-130 and PKU-MMD datasets. Experimental results show that SkelAtom achieves state-of-the-art performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.