Hearing Has a Computable Ceiling
Abstract
Spoken agents are built and evaluated on three assumptions: that the voice carries information the words do not; that a model's loss with and without audio measures what the audio adds; and that speech synthesized from a label can stand in for a person. We prove a ceiling on what hearing can add and test all three assumptions against it. Let be the information, in nats, that the voice carries about the speaker's hidden state beyond her words: no policy gains more than from hearing in one decision, and a turn-by-turn version bounds a whole conversation. Both bounds have sharp constants and are machine-checked in Lean 4. Because is a difference of two log-losses, a system can estimate it from its own conversations, given labels of the user's state and no transcript-reading judge. On television dialogue, the words already carry to percent of what the voice says about the speaker's emotion. The with-and-without comparison overstates what audio adds by the text model's shortfall, six-fold for a bag-of-words text model, and audio from the wrong utterance reproduces most of two off-the-shelf models' apparent gain. Speech synthesized from an emotion label, using instructions we wrote, retains little of what the speaker's own delivery carries. One off-the-shelf model collects a third of the ceiling on dialogue, the other nothing detectable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.