acceptodds
Under review as a conference paper at ICLR 2027

Monoplex: Parallel Speech and Text from a Single Latent

Abstract

A person walking down a street can point, describe the scene aloud and take a note, all at once and from one thought. Machines built on language models say aloud what they write, so their modalities render one content stream rather than working together. We build the parallel alternative, monoplex. In monoplex a frozen vision-language model emits one latent, read at once by its own text head and by a speech head we train. Speech carries a summary, text the detail. With a causal completion stage the speech head reaches 85% of a captioning cascade's content in one pass, first audio at 0.12 s against a streaming cascade's 0.32 s. Trained on triplets that speak the VLM's own descriptions, the speech head learns directly on a speech-distilled codec, and an acoustic codec a third as much on a clean cache. Ranking eight completed speech samples at inference by the text stream and their mutual agreement reaches 98% of the cascade's content on an embedding measure. A coherence judge in that ranking makes two thirds of transcripts coherent, against all of the cascade's, at 96% of its content. On human captions neither stream wrote, the two streams cover more than text alone, and ranking widens the margin.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.