Prior-Centric Support Allocation for Few-Shot Audio Recognition with Audio-Language Models
Abstract
Few-shot audio recognition based on audio-language models requires the collaboration of transferable semantic priors derived from class texts and task-specific evidence from support sets. Nevertheless, existing methods tend to over-emphasize support memories or matching pathways, resulting in the entanglement of sparse support-set evidence with already strong semantic priors and leaving the adaptation mechanism insufficiently explicit. To address this issue, this paper formulates few-shot audio adaptation as a prior-centric support allocation problem, and accordingly proposes a Prior-Centric Support Allocation (PCSA) framework, which treats the support set as a control signal for episode-level inference rather than as a competing memory pathway. The framework consists of two coupled components: Prior-Centric Semantic Scaffold (PriorCore) and Support-Aware Transductive Refinement (SATR). Specifically, PriorCore constructs a stable cross-granular semantic scaffold from coarse and fine priors, while SATR converts support evidence into reliability signals and uses them to regulate transductive refinement over the current episode. Furthermore, this paper designs an explicit reliability-controlled prediction rule that delineates the functional boundaries between semantic priors and support evidence, making the division of labor explicit and analyzable. Extensive experiments on several few-shot audio recognition benchmarks illustrate that our method achieves strong performance under the adopted evaluation protocols, and comprehensive ablation and diagnostic analyses further confirm the effectiveness and robustness of the proposed framework.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.