MotionProbe: Direct LLM Probing for Zero-Shot Skeleton Action Recognition
Abstract
Zero-shot skeleton action recognition commonly represents unseen actions by generating textual descriptions with an LLM and encoding them into semantic prototypes. However, we find that this two-stage process is not neutral: repeated generations under the same prompt produce different recognition outcomes, while auxiliary text encoders impose biases inherited from their own pretraining. The resulting prototypes therefore depend on both how the LLM verbalizes an action and how another model encodes it, showing that semantic construction is a critical part of zero-shot transfer rather than harmless preprocessing. Since the purpose of introducing an LLM is to access its rich action knowledge rather than the descriptions themselves, we propose MotionProbe, which directly reads prompt-conditioned hidden states from a frozen LLM, avoiding the generation variability and encoder bias introduced by the conventional pipeline. Using class-shared motion prompts, MotionProbe probes each action from complementary perspectives and organizes the resulting hidden states into a semantic constellation. These semantic views are then progressively transformed into task-ready prototypes: seen skeletons ground them in discriminative motion geometry, while unlabeled test skeletons further calibrate each candidate constellation and determine how its views should be combined. Extensive experiments on multiple benchmarks demonstrate that MotionProbe attains state-of-the-art performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.