acceptodds
Under review as a conference paper at ICLR 2027

BEYOND LOCAL RECONSTRUCTION: GEOMETRY- PRESERVING OUTPUT ALIGNMENT FOR BINARY LLM QUANTIZATION

Abstract

Recent advances in post-training quantization for large language models (LLMs) have demonstrated that low-bit quantization methods can maintain most of the original model performance. Despite such successes, 1-bit quantization remains particularly challenging. A common strategy in 1-bit quantization is to determine binary weights by matching full-precision parameters, following a weight-driven criterion. However, this objective is not directly aligned with the quantized model's objective, which is to preserve the model's output behavior under the impact of quantization. A natural alternative is to adopt output-driven criteria that minimize discrepancies in model outputs using calibration data. Surprisingly, naive output-driven approaches often perform even worse in the 1-bit regime. In this paper, we show that this failure arises from two fundamental issues: error accumulation across layers and, more critically, anisotropic distortion of the representation space. Based on this observation, we introduce a geometry-preserving output-alignment framework for 1-bit LLM quantization. First, we formulate output alignment as a constrained optimization problem that improves reconstruction while explicitly controlling changes in token similarity, preventing reconstruction gains from being achieved at the expense of representation geometry. Second, we propose a geometry-aware salient-mask selection strategy that identifies quantization configurations according to their ability to preserve token relations, rather than relying solely on reconstruction error. These components enable effective output-driven optimization while retaining the structural information lost by naive alignment. Extensive experiments across multiple LLM families and downstream tasks demonstrate consistent improvements over existing 1-bit PTQ methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.