AVEmbedBench: A Benchmark for Audio-Visual Joint Retrieval in Omni-Embedding Models
Abstract
Recent omni-embedding models support both visual and audio inputs, yet whether they effectively integrate both modalities for retrieval remains unclear. Existing audio-visual benchmarks primarily evaluate generative models through question answering. However, their reliance on textual responses and focus on answer accuracy rather than retrieval rankings limit their applicability to evaluating joint retrieval in embedding models. To address these limitations, we introduce AVEmbedBench, to our knowledge the first benchmark specifically designed to evaluate audio–visual joint retrieval in omni-embedding models. In AVEmbedBench, Each query combines visual and audio-semantic constraints that the target video must jointly satisfy. To discourage unimodal shortcuts, we select clips with weakly correlated visual and audio semantics and pair each target with two hard negatives, one matching only the visual constraint and the other only the audio constraint. To construct the benchmark at scale, we develop an automated pipeline for audio-visual decoupled video filtering, joint query generation, and modality-specific hard negative mining. Evaluations of recent omni-embedding models reveal a consistent visual bias: models favor visual similarity despite audio mismatch, underscoring the gap between multimodal input support and effective joint retrieval, and motivating more balanced audio–visual integration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.