LoopQ: Quantization for Looped Language Models
Abstract
Looped language models (LoopLMs) increase effective computational depth by repeatedly applying shared Transformer blocks without increasing the parameter count. However, this reuse makes them particularly sensitive to post-training quantization (PTQ). We systematically study PTQ in LoopLMs and identify three sources of degradation: distribution shift across roles, state reuse across loop transitions, and recursive error accumulation. We propose LoopQ, a loop-aware PTQ framework that preserves a shared quantized backbone while introducing lightweight adaptations. LoopQ combines loop-aware activation scaling and selective transformations to accommodate changing hidden-state distributions, cross-loop state alignment to stabilize transitions, and trajectory-aware calibration to limit accumulated error. Under W4A4 quantization, LoopQ achieves a 67.8% average relative improvement in five-task mean accuracy across four backbones and an average relative perplexity reduction of at least 75.2% across eight backbone–dataset pairs, compared with the strongest static PTQ baseline for each comparison.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.