acceptodds
Under review as a conference paper at ICLR 2027

Hearing Has a Computable Ceiling

Abstract

Spoken agents are built and evaluated on three assumptions: that the voice carries information the words do not; that a model's loss with and without audio measures what the audio adds; and that speech synthesized from a label can stand in for a person. We prove a ceiling on what hearing can add and test all three assumptions against it. Let be the information, in nats, that the voice carries about the speaker's hidden state beyond her words: no policy gains more than from hearing in one decision, and a turn-by-turn version bounds a whole conversation. Both bounds have sharp constants and are machine-checked in Lean 4. Because is a difference of two log-losses, a system can estimate it from its own conversations, given labels of the user's state and no transcript-reading judge. On television dialogue, the words already carry to percent of what the voice says about the speaker's emotion. The with-and-without comparison overstates what audio adds by the text model's shortfall, six-fold for a bag-of-words text model, and audio from the wrong utterance reproduces most of two off-the-shelf models' apparent gain. Speech synthesized from an emotion label, using instructions we wrote, retains little of what the speaker's own delivery carries. One off-the-shelf model collects a third of the ceiling on dialogue, the other nothing detectable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.