Do Multi-Turn Conversational Capabilities Carry Over from Text to Speech? A Parallel Multilingual Benchmark
Abstract
Successful multi-turn voice interaction requires models to retain information from earlier turns, follow persistent instructions, and track revisions to conversational state. Yet these abilities are often evaluated from transcripts, leaving open whether successes demonstrated from text carry over when the same user turns are presented as speech, and whether this transfer is consistent across languages. We introduce a strictly parallel multilingual extension of Audio MultiChallenge to study this question under controlled conditions. The benchmark contains 354 conversations and 1,404 rubric items across ten languages, with conversational content, turn structure, cross-turn dependencies, and evaluation criteria aligned across transcript and speech conditions. Across six speech-language models, the shift from reference transcripts to direct speech input is strongly model- and language-dependent: models with similar transcript performance can diverge substantially under speech, and their relative strengths across languages can change. Paired rubric-level analysis further shows that aggregate transcript–speech gaps can obscure which conversational requirements are actually preserved, because losses of transcript-demonstrated successes can be offset by gains on other items. We then diagnose recoverable speech degradation by selectively replacing preceding user turns, the current user turn, or both with transcripts generated by the evaluated model. For the two models with the largest transcript–speech gaps, replacing preceding user turns yields a significantly larger macro-level improvement than replacing only the current turn, and the same numerical ordering holds across all ten languages; in contrast, the strongest direct-speech model gains little from either intervention. These results show that transcript performance is an unreliable proxy for multilingual spoken conversational performance and, for models most affected by speech, identify the representation of preceding spoken context as an important source of recoverable degradation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.