acceptodds
Under review as a conference paper at ICLR 2027

SCoPE: Sparse Cross-modal Prior Exchange for Training-Free Audio-Visual Event Perception

Abstract

Audio-Visual Event Perception (AVEP) identifies which events occur, when they occur, and whether they are audible, visible, or both. This distinction matters for sound sources outside the camera's field of view or visible soundless objects. Frozen text-aligned encoders enable training-free recognition with flexible event vocabularies, but events often co-occur and are selected by thresholding scores. When a related incorrect event label scores as high as a correct one, no single threshold can separate them. We analyze this false co-activation (FCA) through two types, shadows and rivals. To address this, we introduce SCoPE, a framework for Sparse Cross-modal Prior Exchange. SCoPE first makes event labels compete to explain the evidence within each modality. Across modalities, prior exchange adjusts event selection costs. Then, target re-selection fits each modality's own embedding under these costs. SCoPE improves event prediction across diverse AVEP benchmarks using the same hyperparameter settings. SCoPE also supports causal streaming for applications that need continuous predictions from incoming audio and video.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.