TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
Abstract
To tackle the high computational cost of large foundation models, post-training activation-aware quantization methods have been widely used. However, because these methods rely heavily on calibration data, calibration–deployment mismatch may arise on unseen downstream tasks. We propose test-time quantization (TTQ), which constructs prompt-specific low-bit weights online from the current prompt and reuses them during decoding. TTQ removes the need for a predetermined calibration dataset when constructing these weights and targets decode-time weight traffic and integer execution rather than resident model-memory reduction. Experiments across language, vision-language, and vision-language-action models show improved low-bit accuracy over offline baselines in several settings while retaining most of AWQ’s decode throughput at a measurable prefill cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.