CAR-Voice-Bench: Benchmark for In-Car Full-Duplex Voice Assistants
Abstract
In-car voice assistants are integral to modern vehicle control, navigation, and telematics. During natural interactions, users routinely exhibit dynamic conversational behaviors, such as concurrent speech, barge-in interruptions, and abrupt goal modifications. Despite their prevalence, existing automotive benchmarks fail to jointly evaluate full-duplex spoken dialogue, executable tool use, and rigorous outcome verification. To address this deficiency, we propose CAR-Voice-Bench, a unified evaluation framework integrating domain-specific tasks, full-duplex interaction, and a replayable testbed. The benchmark comprises 300 curated tasks covering six practical challenges: driving safety, turn management, tool unavailability, disambiguation, simple, and complex. We systematically evaluate three representative models under both standard and dialectal speech conditions with injected background noise, independently testing each configuration three times. Among them, OpenAI's gpt-realtime-2.1 achieves the best overall performance, attaining a pass@1 reward of 51.61%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.