acceptodds
Under review as a conference paper at ICLR 2027

VocalCap: A Speaker-Level Benchmark for Fine-Grained Speech Captioning

Abstract

Understanding speech beyond its words requires describing vocal attributes and associating them with the speakers who produced them. Existing audio-language benchmarks provide limited evaluation of this joint requirement in multi-speaker recordings. We introduce VocalCap, a benchmark for speaker-grounded speech captioning that links fine-grained vocal description with speaker coverage, tempo- ral localization, and attribute-speaker binding. VocalCap comprises 1,000 human- annotated single-speaker clips from film and television and 1,000 controlled mix- tures containing two to six speakers. Both tracks share a 13-field schema spanning perceived speaker characteristics, prosody, articulation, affect, and speaking style. Our evaluation combines field-wise comparison against human references with deterministic temporal matching, reporting coverage, timeline accuracy, and bind- ing for multi-speaker outputs. We evaluate 14 systems on the single-speaker track and 12 on the multi-speaker track. Qwen3.5-Omni-Plus leads both, achieving a normalized single-speaker score of 72.2 and a multi-speaker score of 69.7 on their respective 0–100 scales. Cross-track rank changes show that strong single-speaker captioning does not ensure reliable multi-speaker performance. Scores for affect and style are lower than those for gender presentation and accent, although these subjective fields also have lower annotation agreement. These findings highlight the need to assess caption quality together with speaker coverage and temporal grounding. We release the annotations, source pointers, construction metadata, and evaluation code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.