Phonological Interference in Multilingual Speech Models
Abstract
Phoneme-level speech models recognize or generate speech as a sequence of phonemes, the smallest sound units that distinguish words. These models enable fine-grained pronunciation control and understanding, yet they often fail on input that does not match any single training language, such as speech alternating between two languages, known as code-switching, or speech in low-resource languages absent from training. We identify a systematic failure mode behind this, *phonological interference*: models assume the input is in a single language and impose its phonology, overriding local phoneme-level decisions that conflict with it. We measure interference by how often a model retains phonemes that one language has but the other lacks. On code-switched input, two phone recognizers (speech-to-phoneme models) and a phoneme-conditioned text-to-speech model lose 32%-79% of these phonemes, but lose far fewer of the phonemes both languages share. On unseen languages, recognizers impose the phonology of the training language they assign to the speech: the more confident the assignment, the more they lose phonemes the unseen language has but the assigned language lacks. We probe the models' language estimate from their internal activations, and trace interference to a low dimensional subspace. On monolingual speech, steering this subspace toward a different language makes the model lose the phonemes that only the original language uses and produce those that only the target language has. We introduce *windowed language estimation* (WLE), an inference time repair that replaces the model's language estimate in this subspace with one computed from a short window around each position. On code-switched input, WLE removes 34%-58% of the interference in all three models, and in the phone recognizers it leaves monolingual performance essentially unchanged.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.