Cue-to-Context Subsumption Learning for Audio-Visual Generalized Zero-Shot Learning
Abstract
Audio-Visual Generalized Zero-Shot Learning (AV-GZSL) aims to recognize seen and unseen events by aligning audio-visual samples with textual categories. Although multimodal inputs provide rich information, the representations of full samples are often affected by changing backgrounds and surrounding contexts, making shared event cues difficult to capture across scenes. Existing alignment objectives specify which textual category a full sample should match, but do not explicitly indicate which audio and visual cues should support that match. We observe that recognizable localized cues can recur across different contexts, while each full sample presents these cues under a particular context. Based on this observation, we propose Cue-to-Context Subsumption Learning (CCSL). CCSL first extracts localized visual regions and auditory segments as box inputs. It then maps the box and full inputs into Gaussian distributions using a shared distribution mapper. A directional subsumption objective relates their means and variances, guiding the full representation with localized event cues while retaining information from the complete context. Experiments on VGGSound-GZSL, UCF-GZSL, and ActivityNet-GZSL demonstrate consistent improvements over strong baselines. CCSL improves HM by 5.75 to 7.84 percentage points over the strongest baselines across the three benchmarks and achieves 30.63% unseen accuracy on VGGSound-GZSL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.