Granularity-Aware Evidence Aggregation and Calibration for Training-Free Zero-Shot Multi-Label Recognition
Abstract
CLIP provides strong zero-shot recognition capabilities, but its global representation may overlook small or spatially localized objects in multi-label images. Local crops can recover such evidence, yet may also amplify responses to absent classes, including semantically related alternatives. We introduce GRADE (Granularity-Aware Regional Aggregation and Discriminative Evidence Calibration), a training-free framework for zero-shot multi-label image recognition. GRADE aggregates class-specific regional responses across multiple context granularities, using class-wise global-score percentiles to guide aggregation. When local evidence dominates, comparison with semantically neighboring classes downweights ambiguous local responses before global-local fusion. Experiments on COCO and VG-256 demonstrate competitive performance against existing training-free methods, with consistent mAP improvements over global CLIP across both benchmarks. Ablation studies further confirm the effectiveness of aggregating local evidence across multiple regions and context granularities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.