Real-Time Audio Tool Bench: Evaluating Tool Use in Multi-Turn Voice Interaction
Abstract
Voice-native models now call tools directly from speech, yet how reliably they do so in live, multi-turn conversation is largely unmeasured. Real-time voice raises problems that text tool-calling benchmarks cannot expose: arguments must be heard correctly, details arrive over several turns, and users interrupt, sometimes while a call is still pending. We introduce Real-Time Audio Tool Bench, a benchmark of 1040 tasks across 7 domains and 44 tools, each run live over a provider’s streaming API. Three suites cover explicit multi-turn requests (Reactive), implicit intent (Proactive), and mid-conversation changes of mind (Interruption). Without an LLM judge, deterministic checks grade which tools are called, with what arguments, and when; each failure is attributed to perception (e.g., a misheard name or number) or to reasoning (e.g., a missing, premature, or unexpected call). Across six models from four providers, the best reaches an overall score of 38.7%, and each suite exposes a different bottleneck: text transcripts in place of audio raise Reactive pass rates by 30 to 65 percentage points; models act on nearly all Strong-intent Proactive tasks but complete at most 21% of them; and when a pending call’s result never returns, no model recovers on more than 10.3% of tasks. Static test-time harnesses raise the overall score by at most 5.7 points, and each lowers accuracy on at least one Proactive band or Interruption phase. We release the benchmark, the evaluation framework, and all evaluation traces.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.