Bridge-Duplex: Bridging Semantic and Generative Streams for Full-Duplex Speech-to-Speech Dialogue
Abstract
Full-duplex speech-to-speech (S2S) dialogue requires dialogue-level semantic processing to remain coordinated with continuous frame-wise generation as conversational evidence evolves. Existing systems increasingly make semantic computation explicit, but how semantic state should persist, inform ongoing generation, and be updated throughout dialogue remains underexplored. We propose Bridge-Duplex, a persistent semantic–generative dual-stream framework for full-duplex S2S dialogue. Bridge attention connects the two streams on a shared causal timeline, exposing persistent semantic memory to ongoing generation. Dialogue next-latent prediction further regularizes semantic state evolution by predicting each subsequent semantic state from the current user–system state pair and next-token evidence during training. Bridge-Duplex improves multi-turn dialogue quality and context consistency over the underlying full-duplex models. Selective bridge attention preserves most of the multi-turn semantic performance of dense bridge attention while improving inference efficiency through selective semantic memory access. Audio samples are available at https://anonymous77077.github.io/Bridge-Duplex_Demo/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.