acceptodds
Under review as a conference paper at ICLR 2027

Shared or Independent? A Controlled Study of Transform Search in LLM Quantization

Abstract

Low-bit quantization reduces model storage but introduces approximation error. Rotations can reduce this error by changing the coordinates in which weights are quantized. In attention layers, query, key, and value projections can share an input rotation or use independent rotations. Although independent rotations offer greater flexibility, we show that temporarily sharing Q and K during fitting improves the final independent models on OPT and Llama under matched controls. We study this effect in QKV-only ternary post-training quantization. Our construction copies shared rotations into independent variables without changing the quantized model at the transition, then continues fitting independently. This enables a comparison with independent fitting from the outset while holding the final transform family fixed. We evaluate OPT-1.3B, Llama-3.2-3B, Qwen3-4B-Base, and Ouro-1.4B with matched initializations and restart rules. Temporary sharing lowers negative log-likelihood (NLL) on C4 and WikiText-2 in 19 of 24 model–initialization–corpus comparisons, with 18 improvements remaining significant after correction for multiple comparisons. Additional fitting-budget controls retain the gains on OPT and Llama. Across HellaSwag, ARC-Challenge, and PIQA, normalized accuracy improves in 27 of 36 comparisons, including ten significant improvements after correction. On Llama-3.2-3B, HellaSwag accuracy increases by 1.20 and 0.89 percentage points at the two random initializations. Benefits depend on initialization, sharing duration, and evaluation corpus. Lower reconstruction error can still accompany higher NLL even after accounting for model inputs. Interventions on the fitted models further identify the contributions of Q, K, and layer depth to these performance differences. These results establish temporary sharing as a useful fitting strategy in the studied setting and show why reconstruction error, language-model loss, and downstream accuracy must be distinguished when evaluating quantized models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.