CPU-NPU Cooperative LM-head Quantization for Efficient On-Device Language Generation
Abstract
On-device language generation is constrained by memory capacity and bandwidth. Although low-bit quantization shrinks Transformer decoders, large vocabularies leave the language-model head (LM-head) costly to store and execute on NPUs. However, existing methods either limit compression to prevent quality loss or rely on retrieval-based acceleration that imposes heavy CPU overhead. Efficient edge inference therefore requires reducing both LM-head storage and decoding cost without restricting the vocabulary. We introduce Q2F, which combines NPU-compatible LM-head compression with lightweight CPU refinement using existing embeddings. By exploiting the ability of low-bit predictions to retain promising candidates, Q2F supports greedy and stochastic decoding at 2.2 bits per LM-head weight. Across three models with FP16 decoders, the four-task average accuracy closely matches FP16 baselines, staying within 0.32 percentage points. On Snapdragon 8 Elite Gen 5, Q2F reduces LM-head weight storage by 73% over an INT8 baseline for Llama3.2-1B with a 4-bit decoder, improving end-to-end generation throughput by 19.6% with little CPU refinement overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.