LoAD: Low-Rank Distributional Alignment and Test-Time Decision Adaptation for Zero-Shot Skeleton Action Recognition
Abstract
Generative methods for zero-shot skeleton action recognition mostly follow an align-then-classify paradigm, where a cross-modal autoencoder aligns skeleton sequences with class semantics in a shared latent space, upon which a classifier is built. This paper revisits both phases. On the alignment side, class-conditional latent distributions are typically modeled as diagonal Gaussians, leaving the correlation structure within each distribution unexploited. On the decision side, different decisions usually read the same generative score, whereas they in fact impose different requirements on the evidence. Accordingly, we propose LoAD, which improves this paradigm in both alignment depth and decision structure. First, Low-Rank Covariance Alignment (LRCA) extends the latent distribution to a diagonal-plus-low-rank structure, and jointly aligns the means, variances, and low-rank correlation factors of both modalities during training. Second, Test-time Decision Adaptation (TDA) refines how decisions are made at test time, with two components. Decision-Decoupled Inference (DDI) constructs two complementary per-class evidences and reads them in a decoupled manner, where class ranking in the zero-shot setting fuses a flow-matching velocity error with a deterministic reconstruction error, while seen/unseen discrimination in the generalized setting relies on the former only. Test-time Shrinkage Calibration (TSC) further adapts the class latents with statistics from unlabeled test batches, modulated by the number of seen classes and requiring no gradient updates. On standard splits of NTU-60, NTU-120, and PKU-MMD, LoAD achieves competitive performance under both zero-shot and generalized zero-shot settings, and systematic diagnostic studies verify the effectiveness of each design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.