acceptodds
Under review as a conference paper at ICLR 2027

Quantization Modulates Batch-Context Non-Determinism in Greedy LLM Inference

Abstract

Quantization reduces the memory and compute cost of LLM inference, and requests are usually batched to serve many of them at once. However, recent studies have shown that batching causes non-determinism in LLM inference despite the use of greedy decoding. We study whether quantization changes this batch-context non-determinism, holding prompts, decoding settings and checkpoints fixed. For that we evaluate primarily Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct across seven benchmarks and six serving contexts using BF16, GPTQ, FP8 weight-and-activation quantization, and FP8 KV-cache quantization. We observe answer disagreement across every pair of our serving contexts. FP8 weight-and-activation quantization roughly doubles how often an example changes its answer, from 8.7% to 17.2% for Qwen2.5-7B-Instruct and from 13.9% to 25.3% for Llama-3.1-8B-Instruct, while weight-only GPTQ stays close to BF16. Quantizing only the KV cache also raises the rate, and smaller evaluations on five further models show the same pattern. We trace the cause to the order in which floating-point reductions are carried out, which depends on batch composition. When we use batch-invariant kernels, the non-determinism effect vanishes for BF16 and FP8 at a throughput loss, but it remains for GPTQ.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.