AQUA: Adaptive Querying and Alignment for Invariant Prototype Learning in Few-shot Action Recognition
Abstract
Few-shot action recognition (FSAR) aims to recognize novel actions from only a few labeled videos. However, existing methods often struggle to construct reliable and class-consistent prototypes from noisy support sets. We propose AQUA, an Adaptive QUerying and Alignment framework for invariant prototype learning in FSAR. AQUA introduces Invariant Prototype Querying (InvQ), in which learnable queries selectively aggregate class-consistent information across support examples via commonality-biased cross-attention, thereby suppressing sample-level variations. To semantically regularize the raw visual representations, Commonality Alignment (CMAlign) provides an auxiliary training objective that aligns the support representations from InvQ and query representations with class-wise textual anchors. Furthermore, Commonality Adaptation (CMAdapt) dynamically refines textual prototypes through consistency-gated visual conditioning, allowing reliable visual evidence to complement textual semantics. Extensive experiments demonstrate that AQUA consistently improves recognition performance across diverse few-shot action recognition settings, highlighting the effectiveness and generalizability of the proposed framework.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.