Learning Visual Relations with Adaptive Semantic Calibration for Few-Shot Action Recognition
Abstract
Few-shot action recognition (FSAR) aims to recognize novel action categories from only a few labeled videos. Recent approaches have benefited from set-based video matching and semantic knowledge derived from pretrained vision-language models (VLMs). However, semantic knowledge remains unevenly utilized in FSAR: it is underused in support-query relation learning, insufficiently adapted to video-level recognition, and combined with visual evidence under globally shared supervision that ignores query-specific differences. To address these limitations, we propose Semantic-Adaptive Relation Fusion (SARF), which couples semantic-guided visual relation learning with semantic response recalibration and instance-adaptive semantic–visual supervision. Specifically, the Text-Guided Feature Alignment (TFA) module incorporates textual knowledge into support-query relation modeling to strengthen discriminative local correspondences between videos. The Semantic Logits Recalibration (SLR) module learns residual corrections for VLM-derived semantic logits, adapting semantic responses to the target FSAR task and alleviating confusion among semantically related categories. Finally, the Semantic–Visual Fusion (SVF) module estimates instance-dependent weights from visual relation representations and constructs adaptive semantic–visual predictions that provide query-specific auxiliary supervision during training. Extensive experiments on multiple FSAR benchmarks demonstrate that SARF achieves state-of-the-art performance across most evaluated settings, validating the effectiveness of jointly learning fine-grained visual relations and adaptively coordinating semantic and visual evidence under few-shot supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.