acceptodds
Under review as a conference paper at ICLR 2027

TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly

Abstract

To tackle the high computational cost of large foundation models, post-training activation-aware quantization methods have been widely used. However, because these methods rely heavily on calibration data, calibration–deployment mismatch may arise on unseen downstream tasks. We propose test-time quantization (TTQ), which constructs prompt-specific low-bit weights online from the current prompt and reuses them during decoding. TTQ removes the need for a predetermined calibration dataset when constructing these weights and targets decode-time weight traffic and integer execution rather than resident model-memory reduction. Experiments across language, vision-language, and vision-language-action models show improved low-bit accuracy over offline baselines in several settings while retaining most of AWQ’s decode throughput at a measurable prefill cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.