acceptodds
Under review as a conference paper at ICLR 2027

RIVA: An Open Recipe for Agentic Full-Duplex Speech Models

Abstract

Full-duplex speech language models let users interact with an AI assistant as they would with another person. Existing open models, however, struggle to combine native interaction, in which the model itself decides when to listen, speak, or yield, with the language and agentic capabilities of a strong large language model (LLM). Tool use is also typically handled outside the conversational stream, through separate action channels or backend models. We introduce Riva, an open full-duplex speech model in which a pretrained LLM listens, speaks, and acts within a single stream. Every 480 ms, the model processes incoming speech, decides whether to listen or speak, generates text that a lightweight talker renders as speech, and can emit tool calls in its native format while the conversation continues. Three components make this possible under real speech: a pre-training recipe that reduces the speech-text modality gap; Riva-Dialog, approximately 500K skilled, speech-native conversations spanning 13K hours, complemented by Riva-Entity for named entities and structured identifiers; and reinforcement learning over short decision windows, combining GRPO for interaction-level decisions with on-policy self-distillation for exact tool arguments. Riva is competitive with or outperforms prior open full-duplex models across six benchmarks, including Audio MultiChallenge (18.4 vs. 12.8 APR) and Full-Duplex-Bench v3 (42.0 vs. 35.5 P@1), and achieves an average P@1 of 28.6 on multi-minute, multi-tool customer-service interactions in τ²-Voice.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.