acceptodds
Under review as a conference paper at ICLR 2027

SCOPE-Q: ONE-CHECKPOINT QUANTIZATION-AWARE TRAINING FOR CROSS-FRAMEWORK LLM DEPLOYMENT

Abstract

Large language models (LLMs) are often deployed through several frameworks, such as OpenVINO, TensorRT, Core ML, and AIMET, each of which quantizes the model with its own quantization configuration. Most LLM quantization methods optimize a model for a single, fixed quantizer. We find that this assumption is costly in practice: at the same 4-bit precision, the MMLU accuracy of one checkpoint varies by 12–15 points across these four frameworks, mainly because of differences in their quantization configurations. In this paper, we propose SCOPEQ, a unified quantization-aware training method that produces a single checkpoint for multiple deployment frameworks. SCOPE-Q learns a low-rank update to the frozen pretrained weights while sampling one target quantization configuration at each training step, and distills knowledge from the full-precision model. After training, the update is merged into the weights, yielding a standard floating-point checkpoint that each framework quantizes without modification. Experiments on four LLMs ranging from 0.5B to 4.6B parameters show that a single SCOPE-Q checkpoint improves worst-case accuracy across frameworks by 7.5–8.4 points over round-to-nearest quantization and reduces the accuracy gap between frameworks by 44–69%, while training only 1.1–3.6% of the parameters.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.