Sparse Continuous Representation for Unpaired Speech–Text Learning
Abstract
Learning from unpaired speech and text corpora can substantially improve sample efficiency and accessibility of spoken language systems. However, existing approaches for unpaired speech–text learning face a **capacity–compatibility tradeoff**: Dense, continuous representations preserve rich acoustic information but differ significantly with text in representational structures, resulting in low compatibility; Discrete speech units improve compatibility at the cost of information loss. In this work, we propose **sparse continuous representations (SCRs)** as an alternative that preserves rich information while maintaining compatibility. To validate our approach, we consider the task of unsupervised speech recognition (UASR) and synthesis (UTTS) as testbeds, both canonical examples of unpaired speech-text learning. Across seven languages for UASR, our SCR-based systems, SparPUSM and SparCipher, consistently outperforms the previous state-of-the-art systems, yielding an average character error rate reduction of 8-17%, with up to 17% on LibriSpeech. For UTTS, SCR-based approach also outperforms previous approaches, demonstrating the benefit beyond UASR. Code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.