acceptodds
Under review as a conference paper at ICLR 2027

Dy-VoiceBench: Stateful Worlds and Living Users for Evaluating Dynamic Spoken Dialogue Agents

Abstract

Conversational agents are increasingly expected to complete tasks through interaction rather than generate isolated responses. However, many existing evaluations still treat dialogue as static prediction, where future turns, tool contexts and user behaviors are fixed and the model is not fully accountable for the consequences of its own trajectory. This is insufficient for spoken task completion, where model utterances and actions may change the external world, the user's ability to cooperate, and the episode-local knowledge available in later turns. We introduce Dy-VoiceBench, a benchmark for evaluating dynamic spoken dialogue agents through executable closed-loop episodes. Each episode specifies hidden states, visible observations, model-side and user-side actions, environment feedback, transition rules, process rubrics, and terminal success criteria. Rather than matching a reference transcript, Dy-VoiceBench scores the state trajectory induced by the evaluated model. It covers three complementary dynamics: world-state dynamics for executable task completion and user-action coordination, user-state dynamics for task-consequential spoken empathy and cooperation maintenance, and knowledge-state dynamics for acquiring, preserving, revising, and applying episode-local knowledge. Dy-VoiceBench contains 614 scenarios and 1,632 executable world-state action instances, providing a diagnosable testbed for state-centric spoken dialogue evaluation. Demo is available at https://dyvoicebench.github.io/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.