acceptodds
Under review as a conference paper at ICLR 2027

Revision-Bound Grounding for Stateful Tool Use in Open-Source Full-Duplex Voice Agents

Abstract

Full-duplex voice agents can listen, speak, reason, and act concurrently, but this creates a state-consistency problem: user intent may change while earlier reasoning or tool execution remains in flight, allowing stale evidence or actions to affect the interaction. We introduce *revision-bound grounding*, a causal interface that binds asynchronous evidence, action proposals, authorizations, and tool outcomes to their originating dialogue revision and action trace, admitting them only while they remain authoritative. We instantiate this interface in MoshiAgent, which decouples continuous speech from asynchronous reasoning and tool execution while maintaining explicit dialogue, action, and observed world state. We also construct a tool-grounded duplex corpus of conversations totaling hours and train MoshiAgent-Speech to incorporate delayed evidence during ongoing interaction. Across six audio-agent benchmark families and ToolRaceBench, a diagnostic benchmark for asynchronous state consistency, revision-bound execution with fixed MoshiAgent-Speech improves Full-Duplex-Bench v3 argument accuracy from to and strict from to , while increasing ToolRaceBench from to . The complete system achieves Macro on -Voice, surpassing GPT-Realtime-2.1 mini. These results show that reliable full-duplex tool use requires explicit causal alignment between evolving user intent, external actions, and observed outcomes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.