Learning Compressed Audio-Visual Representations via Conditional Causal Distillation
Abstract
Audio-visual distillation compresses large multimodal teachers by transferring predictions or representations to compact students. Although existing objectives employ different knowledge carriers, they generally assume that teacher targets provide reliable supervision. However, teachers learned from observational data may entangle task semantics with latent multimodal priors arising from scene context, modality quality, and cross-modal correspondence. Direct imitation can thus preserve both useful knowledge and context-dependent response patterns, especially under limited student capacity. We present Causal Audio-Visual Distillation (CAVD), a teacher-student framework that standardizes supervision before compression. CAVD captures recurrent multimodal context patterns with a compact latent state set and models teacher responses conditioned on each state. It then standardizes these responses over a shared reference distribution, producing supervision that is less dependent on individual sample context composition. The student learns both standardized predictions and state-conditioned response variations without auxiliary annotations or modality-specific bias modules. Operating in latent representation space, CAVD provides a unified formulation for mutually exclusive, multi-label, and dense outputs. Experiments on temporal localization, video parsing, and sounding-object segmentation demonstrate consistent improvements over distillation baselines, with compact students approaching teacher performance while using approximately 80-95% fewer parameters and FLOPs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.