SyncAudio: Parallel and Synchronized Text-Speech Decoding for Spoken Dialogue
Abstract
Extending text-based LLMs to spoken dialogue requires integrating speech generation without sacrificing the linguistic and knowledge capabilities of the underlying model. We study this problem in a shared-backbone architecture with parallel text and speech generation, where the two modalities are decoded through separate pathways but interact through a common language-model context. A central challenge is how the speech pathway interacts with the semantic structure inherited from the text LLM. Such interaction can introduce degradation in two complementary ways. First, temporal mismatch between text and speech can mix states from different stages of the response, weakening the semantic coherence of the shared context. Second, the speech pathway depends on backbone representations whose cross-modal relevance varies across depth, so conditioning on a single layer may either miss complementary information or emphasize representations that interfere with semantic preservation. We introduce SyncAudio, which addresses these two sources of cross-modal interference through dynamic text-speech synchronization and sparse multi-depth aggregation. For temporal coordination, SyncAudio uses an audio-guided synchronization mechanism to determine when the evolving speech state should be committed and fed back into the shared context. For representational interaction, it sparsely selects and aggregates representative hidden states from different backbone depths to provide richer conditioning for speech generation, while leaving the standard text-decoding pathway unchanged. Experiments show that dynamic synchronization substantially improves text-speech alignment while preserving natural and intelligible speech generation. More importantly, SyncAudio consistently outperforms an otherwise matched fixed-synchronization baseline on spoken question answering, suggesting better preservation of the semantic and knowledge capabilities inherited from the underlying text LLM. Further ablations show that sparse multi-depth representation aggregation provides more effective conditioning for the speech pathway than relying on the final backbone layer alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.