MAD-SP: Multi-Agent Dynamic Semantic Prompting Framework for Few-Shot Image Recognition
Abstract
Few-shot learning (FSL) recognizes novel classes from few labeled examples, but scarce support data yield unstable visual representations. Multimodal methods alleviate this limitation by incorporating textual semantics as prompts, but LLM-generated captions may be overly verbose and redundant. To address these issues, we propose Multi-Agent Dynamic Semantic Prompting (MAD-SP), which improves both the construction and utilization of semantic prompts. First, Multi-Agent Caption Refinement (MACR) coordinates generator, critic, and refiner agents under CLIP-based alignment and discriminative scoring to produce concise, visually grounded captions, thereby providing higher-quality semantic prompts for visual recognition. Building on these refined prompts, Dynamic Prompt Injection (DPI) enables their instance-adaptive utilization by employing a lightweight learnable decision network to estimate the compatibility between each semantic prompt and visual features at different blocks, thereby selecting an appropriate injection layer for each image. To ensure reliable alignment between semantic prompts and layer-specific visual representations during varying-depth injection, Contrastive Modality Alignment (CMA) further supports DPI by introducing a bottleneck alignment adapter after the text encoder and optimizing it with the InfoNCE loss. Together, MACR improves the quality of semantic prompts, while DPI and CMA enable their effective and adaptive integration into the visual encoder. Extensive experiments show that MAD-SP achieves state-of-the-art results in all eight settings on four standard benchmarks and in seven of eight settings across four cross-domain benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.