Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace
Abstract
Recent interpretability work shows that language models maintain a small, privileged set of verbalizable representations in their middle layers—readable as words, reused across internal computation, and broadcast to later processing—that behaves like a global workspace, so far characterized for text inputs. We show that an audio language model forms such representations from sound. Reading an off-the-shelf Qwen3-Omni model with a logit lens at the audio-token positions, which precede and cannot see the question, we find answer-relevant concepts readable as words from the early and middle layers on, before any token is emitted. A waveform-swap control fixes the question and options and swaps only the audio (the real clip, a mismatched clip, or silence): the clip-level readout is best with the real clip, weaker with a mismatched clip, and at chance with silence, so it is driven by the sound rather than by the lens's preference for some option tokens. Traced across depth, this sound-driven signal is null at the input, becomes significant about a tenth of the way up, and persists to the output. Patching the audio-position activations into a run whose audio was removed restores the answer from the early and middle bands but not from the late layers, so the audio content is used, and committed before the late layers. The readable content is conceptual rather than a re-transcription: one concept surfaces in several languages and scripts, and a speaker's emotion is read at the audio positions although the model's own transcript and caption usually do not name it. On a clip about a president forced to resign, Watergate and scandal are followed by the answer, Nixon, at positions that cannot see the printed options. In this model, the verbalizable representations that text models form thus also form from sound.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.