RIVA: An Open Recipe for Agentic Full-Duplex Speech Models
Abstract
Full-duplex speech language models let users interact with an AI assistant as they would with another person. Existing open models, however, struggle to combine native interaction, in which the model itself decides when to listen, speak, or yield, with the language and agentic capabilities of a strong large language model (LLM). Tool use is also typically handled outside the conversational stream, through separate action channels or backend models. We introduce Riva, an open full-duplex speech model in which a pretrained LLM listens, speaks, and acts within a single stream. Every 480 ms, the model processes incoming speech, decides whether to listen or speak, generates text that a lightweight talker renders as speech, and can emit tool calls in its native format while the conversation continues. Three components make this possible under real speech: a pre-training recipe that reduces the speech-text modality gap; Riva-Dialog, approximately 500K skilled, speech-native conversations spanning 13K hours, complemented by Riva-Entity for named entities and structured identifiers; and reinforcement learning over short decision windows, combining GRPO for interaction-level decisions with on-policy self-distillation for exact tool arguments. Riva is competitive with or outperforms prior open full-duplex models across six benchmarks, including Audio MultiChallenge (18.4 vs. 12.8 APR) and Full-Duplex-Bench v3 (42.0 vs. 35.5 P@1), and achieves an average P@1 of 28.6 on multi-minute, multi-tool customer-service interactions in τ²-Voice.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.