Hardware-Aware Language Model Quantization For A Many-Core AI Inference Accelerator
Abstract
Edge accelerators for ML/AI workloads are designed around trade-offs between peak performance, memory capacity and bandwidth, power, and cost, with vendor economics driven by volume rather than per-unit margin. Because deployed silicon cannot be refreshed on the cadence of model releases, gains must increasingly come from the algorithmic side: pushing out the Pareto frontier of accuracy, memory footprint, and throughput within the constraints of hardware already in the field. Given the extensive work on post-training quantization, we choose it as the primary compression lever and treat it as a full inference-stack optimization problem, jointly optimizing accuracy, storage, and throughput rather than minimizing quantization error alone. We account for hardware features such as supported data formats, computational units, and the instruction set architecture, alongside structural constraints imposed by the runtime and compiler. Beyond the formats natively supported by the hardware, we treat the format design space itself — element and scale datatypes, group size, and hierarchical scales — as part of the optimization process, covering standardized block formats such as MXFP and NVFP as well as non-standard constructions. Algorithmically, we evaluate mixed-precision allocation and codebook-based quantization across dense and mixture-of-experts models from 4B to 35B parameters. We share with the wider community insights on what works well on paper and what actually improves the Pareto frontier across accuracy, memory footprint, and throughput for models deployed on edge silicon. We share with the wider community insights and trade-offs on what works well on paper and what actually improves the Pareto frontier across accuracy, memory footprint, and throughput for a model deployed on edge silicon. We demonstrate our hardware-aware quantization methodology on a many-core edge accelerator with over a thousand RISC-V cores, achieving 34%, 67%, and 24% decode speedups over the llama.cpp K-quant of equal storage at 4.3, 3.7 and 3.3 weight-memory reduction over BF16, within one point of its mean accuracy, on Qwen3.5-35B-A3B, Qwen3.5-4B, and Llama-3.1-8B respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.