Monoplex: Parallel Speech and Text from a Single Latent
Abstract
A person walking down a street can point, describe the scene aloud and take a note, all at once and from one thought. Machines built on language models say aloud what they write, so their modalities render one content stream rather than working together. We build the parallel alternative, monoplex. In monoplex a frozen vision-language model emits one latent, read at once by its own text head and by a speech head we train. Speech carries a summary, text the detail. With a causal completion stage the speech head reaches 85% of a captioning cascade's content in one pass, first audio at 0.12 s against a streaming cascade's 0.32 s. Trained on triplets that speak the VLM's own descriptions, the speech head learns directly on a speech-distilled codec, and an acoustic codec a third as much on a clean cache. Ranking eight completed speech samples at inference by the text stream and their mutual agreement reaches 98% of the cascade's content on an embedding measure. A coherence judge in that ranking makes two thirds of transcripts coherent, against all of the cascade's, at 96% of its content. On human captions neither stream wrote, the two streams cover more than text alone, and ranking widens the margin.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.