acceptodds
Under review as a conference paper at ICLR 2027

Single Thread, Real Time: Heterogeneous Quantization for CPU Speech Recognition

Abstract

LM-based speech recognition models are among the most accurate ASR systems today, but they typically require GPUs for real-time inference, so private audio leaves the device and latency is tied to the network. We move them onto device CPUs and propose a memory-traffic model that assigns a quantized format per module. The transformer is weight-bound, so its weights are ternarized; the acoustic stack is activation-bound, so it runs entirely in INT8, which shrinks activation memory 6.2x against FP32. Progressive fake-quantization keeps training from diverging, and AVX2 and NEON intrinsics accelerate both formats. Trained and deployed this way, a 1.5B recognition model is up to 4.32x faster on a single thread than FP16 and 2.9x smaller, and transcribes faster than real time on one thread of either an Apple M4 or an Intel Core i7-13700, for a 0.43% increase in recognition error across ten corpora relative to the FP16 baseline. The allocation transfers to other recognition and generation models with comparable gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.