acceptodds
Under review as a conference paper at ICLR 2027

Cue-to-Context Subsumption Learning for Audio-Visual Generalized Zero-Shot Learning

Abstract

Audio-Visual Generalized Zero-Shot Learning (AV-GZSL) aims to recognize seen and unseen events by aligning audio-visual samples with textual categories. Although multimodal inputs provide rich information, the representations of full samples are often affected by changing backgrounds and surrounding contexts, making shared event cues difficult to capture across scenes. Existing alignment objectives specify which textual category a full sample should match, but do not explicitly indicate which audio and visual cues should support that match. We observe that recognizable localized cues can recur across different contexts, while each full sample presents these cues under a particular context. Based on this observation, we propose Cue-to-Context Subsumption Learning (CCSL). CCSL first extracts localized visual regions and auditory segments as box inputs. It then maps the box and full inputs into Gaussian distributions using a shared distribution mapper. A directional subsumption objective relates their means and variances, guiding the full representation with localized event cues while retaining information from the complete context. Experiments on VGGSound-GZSL, UCF-GZSL, and ActivityNet-GZSL demonstrate consistent improvements over strong baselines. CCSL improves HM by 5.75 to 7.84 percentage points over the strongest baselines across the three benchmarks and achieves 30.63% unseen accuracy on VGGSound-GZSL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.