Pushing the Limits of Sub-10M Multi-Domain Speech Recognition
Abstract
Edge speech recognition requires multi-domain robustness under strict parameter caps, yet standard uniform compression degrades deep acoustic representations. In this paper, we propose AcuFormer, the first sub-10M ASR model designed for robust multi-domain performance within a complete 9.97M parameter envelope through an acoustically informed architecture. Guided by acoustic sensitivity, AcuFormer preserves full encoder depth while reallocating released feed-forward parameters into compact bidirectional GRU decoders that reuse acoustic token embeddings, targeted local recovery modules, and a blank-invariant prior for multi-view selection. Evaluated across all seven domains of the OpenASR benchmark, AcuFormer achieves 9.70% average WER, breaking the 10% error threshold for sub-10M models, yielding a 59.0% relative improvement over the compact state of the art, and outperforming baselines up to 7.8 times its size. Furthermore, real-world hardware profiling on a low-power Raspberry Pi Zero 2 W confirms real-time execution with an RTF of 0.53 and minimal memory residency. Taken together, these results push the limits of sub-10M multi-domain speech recognition by uniting deep acoustic accuracy with extreme on-device efficiency, facilitating private, always-available voice interfaces on ubiquitous edge devices.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.