Speech recognizers carry the surrounding speech rate but use it less as they grow
Abstract
A probe can decode a variable from a network's activations without the network's behavior depending on it. Separating a variable the network carries from one it uses calls for a manipulation few variables permit: the variable moves while the input the decision is about stays fixed. Human speech perception supplies one. When the speech around a reduced function word such as “or” slows, listeners' reports of it fall from 79 to 33 percent though its audio is unchanged. A recognizer that did the same would drop a correctly transcribed word when only the speech around it slowed, a change that word error counts but cannot attribute to the context. We ran this manipulation on sixteen open checkpoints, twelve Whisper models across five sizes and four CTC recognizers. Whisper tiny through medium lower their belief in the frozen word, the log-probability lead of the transcript with the word over the one without it, as the surrounding speech slows. In the three smaller models the fall is roughly a third to half of the listeners' own, and the word drops out of the transcript. Three further function words and the CTC recognizers show the shift as well. Within Whisper the shift declines with size, persisting at 1.5B in two of three checkpoints, while the context's rate stays decodable from the same frozen frames at every size. Erasing that representation does not selectively remove the shift in Whisper and enlarges it in two of the three primary CTC recognizers. Together, this work separates a speech recognizer's representation from its use, which no benchmark that alters only the target, and no probe, can see. Our approach gives interpretability a stimulus-side test of use, complementary to probes and erasure, that costs one item set and no retraining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.