VanuatuASR: Open Multilingual Speech Recognition for Low-Resource Vanuatu Languages
Abstract
Multilingual automatic speech recognition (ASR) remains inaccessible to many languages that lack transcribed speech, digital text, and representation in large-scale pretraining corpora. We present VanuatuASR, to our knowledge the first unified multilingual ASR system for Vanuatu, a highly linguistically diverse region where speech and text resources are severely limited. Building on speech curated across seven Vanuatu languages totaling 65.92 hours, we investigate three practical questions: whether joint multilingual fine-tuning can match language-specific recognizers; whether a shared language model trained on pooled transcripts improves decoding when per-language text is too scarce to support individual models; and whether adding untranscribed speech from additional languages benefits self-supervised representation learning. Our system combines SSL on pooled unlabeled speech, language-balanced joint CTC fine-tuning on four transcribed languages, and shared character -gram LM decoding without language identifiers. Our results show that joint training matches four language-specific models with a single recognizer. The shared LM is the dominant source of improvement, reducing macro WER from 37.37% to 32.12% and remaining effective with as little as 25% of transcribed training data. Expanding the SSL pool with untranscribed speech from three additional languages, however, yields no measurable gain. VanuatuASR outperforms other state-of-the-art baselines under a common evaluation protocol, reducing macro WER by 10.89% relative to the strongest baseline. These results indicate that in severely low-resource settings, transcribing existing recordings and pooling scarce textual resources across languages delivers more value than collecting additional unlabeled speech.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.