Single Thread, Real Time: Heterogeneous Quantization for CPU Speech Recognition
Abstract
LM-based speech recognition models are among the most accurate ASR systems today, but they typically require GPUs for real-time inference, so private audio leaves the device and latency is tied to the network. We move them onto device CPUs and propose a memory-traffic model that assigns a quantized format per module. The transformer is weight-bound, so its weights are ternarized; the acoustic stack is activation-bound, so it runs entirely in INT8, which shrinks activation memory 6.2x against FP32. Progressive fake-quantization keeps training from diverging, and AVX2 and NEON intrinsics accelerate both formats. Trained and deployed this way, a 1.5B recognition model is up to 4.32x faster on a single thread than FP16 and 2.9x smaller, and transcribes faster than real time on one thread of either an Apple M4 or an Intel Core i7-13700, for a 0.43% increase in recognition error across ten corpora relative to the FP16 baseline. The allocation transfers to other recognition and generation models with comparable gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.