VoiceWorld-Bench: Evaluating AI Voice Agents Across Listening, Acting, and Talking in Realistic Environments
Abstract
As large language models evolve into real-time speech-interactive systems, AI voice agents can listen, reason, use tools, and speak across consumer and enterprise workflows. Yet existing benchmarks do not evaluate them end to end: agent benchmarks focus on text- or GUI-driven tool use, while audio benchmarks primarily measure speech understanding or dialogue quality, without requiring grounded action or task completion. To close this gap, we introduce VOICEWORLD-BENCH, a benchmark for AI voice agents in two realistic task environments: a voice-commerce online storefront and a voice-driven coding IDE. Each environment exposes MCP-based tools over persistent backend state, enabling agents to act in the environment and allowing deterministic verifiable judges to measure task completion. VOICEWORLD-BENCH evaluates AI voice agents along four dimensions: task completion, speech robustness, real-time responsiveness, and safety & security. We evaluate three agent designs: audio-native agents, which use a single model for speech interaction, reasoning, and tool calling; transcription-cascaded agents, which combine a dedicated transcription module, a text LLM, and a text-to-speech (TTS) module; and audio-cue-aware agents, which use an audio-capable LLM for paralinguistics-aware speech understanding before text-based tool execution. Our evaluation highlights several emerging findings. First, current AI voice agents are far from saturating VOICEWORLD-BENCH and exhibit distinct limitations across evaluation perspectives. Second, cascaded designs often achieve stronger grounded execution than fully audio-native agents, suggesting that text-based reasoning and tool use remain advantageous when task success depends less on audio-specific cues. Third, audio-cue-aware understanding benefits semantic intent recovery, but can come at the cost of degrading high-precision symbolic recovery. Fourth, safety and capability remain decoupled: agents that complete benign tasks successfully do not necessarily behave safely under adversarial conditions. Together, VOICEWORLD-BENCH and these findings motivate holistic evaluation of AI voice agents, and point to future directions in both agent-harness design and backbone-model adaptation for speech-aware, tool-grounded execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.