acceptodds
Under review as a conference paper at ICLR 2027

Voca: Are Large Audio Language Models Ready for Voice Companionship?

Abstract

Large Audio Language Models (LALMs) have made substantial progress in speech understanding and generation, yet voice companionship asks for more: perceiving how a user is doing and responding with appropriate emotion, and beyond these, offering unrequested support and staying reliable in sensitive situations. Current systems handle the first two on request, but do not take up the last two on their own. We term this condition reactive companionship. Companionship, however, is largely constituted by what is never asked for, and this is exactly what existing LALM benchmarks leave unmeasured. Targeting audio understanding, reasoning, and speech generation, they evaluate capabilities the user must first call upon. To address this gap, we introduce Voca, a benchmark that evaluates voice companionship along four dimensions: User State Understanding, Emotional Interaction, Proactive Care, and Safe Companion Behavior. Across fourteen end-to-end, real-time, and cascaded systems, competence in perception does not carry over. Under a standard system prompt, no system exceeds 48.39 on proactive care, and explicitly prompting models to attend to users' latent needs still leaves thirteen of fourteen below 50. We further propose VocaAgent, an inference-time multi-agent framework that separates cue monitoring from response deliberation and substantially improves proactive care. The demo could be found at https://vocaresearch.github.io/voca/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.