AquaPIC: Boosting On-device LLM Quantization via Prefetching Input-wise Compensation
Abstract
Post-training quantization (PTQ) enables efficient local LLM inference, but aggressive quantization can substantially degrade accuracy. Recovering this lost accuracy typically requires additional correction data, but the gains are constrained by limited device memory budget. Prefetching correction data from second-tier memory can overcome this capacity limit, yet limited transfer bandwidth constrains their size, making compact low-rank compensation particularly suitable for this setting. We introduce input-partition-wise compensation (IPC), which exploits additional second-tier capacity by storing specialized low-rank compensators for different input regions without increasing the amount of data fetched per input. Building on IPC, AquaPIC combines output-error-aligned clustering, shared block-level partitions, and a lightweight residual-based partition estimator to select and prefetch compensators while the preceding block executes. Across two LLMs and two quantization methods, AquaPIC improves zero-shot accuracy under 2-bit quantization by 6.8 percentage points on average over the strongest same-rank baselines, while incurring only 2.27% average latency overhead across discrete GPU and unified memory setups.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.