EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
Abstract
Quantization has emerged as a mainstream approach for deploying Large Language Models (LLMs) on resource-constrained devices, yet reducing precision below 4-bit typically leads to severe performance degradation or prohibitive retraining costs. In this paper, we propose EdgeRazor, a lightweight framework for LLMs via Mixed-Precision Quantization-Aware Distillation. It contains three modules: Structural Quantization with Mixed Precision for fine-grained control over bit-widths, Layer-Adaptive Feature Distillation for dynamically selecting the most informative features for alignment, and Entropy-Aware KL Divergence for balancing forward and reverse KL to handle data from both human annotation and model distillation. With 1.88-bit weights, 8-bit activations, and 8-bit KV cache, EdgeRazor on Qwen3-0.6B outperforms the strongest 2-bit baseline by 11.40 points in average accuracy and the strongest 3-bit baseline by 4.38 points. In terms of training efficiency, EdgeRazor quantizes MobileLLM-350M with 4-10 fewer training tokens than the leading quantization-aware training method. On the deployment side, EdgeRazor achieves consistently higher compression ratios than existing methods across all bit-widths. Exported to the TQ2_0 quantization type supported by llama.cpp, the 1.58-bit Qwen3-0.6B quantized by EdgeRazor reduces the model size from 1.11 GB to 0.19 GB and decodes 15.16 faster than its 16-bit counterpart. Extensive experiments on the MobileLLM and Qwen families confirm the effectiveness and efficiency of EdgeRazor.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.