acceptodds
Under review as a conference paper at ICLR 2027

Attend to Your Own Thoughts: Calibration Matters for 1.58-Bit Reasoning LLMs

Abstract

In this paper, we present ScaleQ-1.58, a ternary post-training quantization (PTQ) framework for reasoning LLMs. Its core insight stems from an empirical finding: although modern LLMs are typically trained to have chain-of-thought reasoning capabilities, in the PTQ regime, even recently proposed differentiable ternarization still leads to performance collapse on challenging mathematics and coding tasks. We reveal that this is caused by conventional calibration schemes that ignore the model's reasoning process. Motivated by this, we present Attend to Your Own Thoughts (AYOT), a new and simple calibration method, which constitutes the core contribution of this paper. Given a pre-trained target LLM, AYOT uses its self-generated reasoning traces and final answers on a proper set of calibration samples, together with the corresponding questions, as context inputs during ternarization. ScaleQ-1.58 is built by combining AYOT with differentiable ternarization, and demonstrates several favorable scaling properties. With only 4M calibration tokens: (1) Qwen3-1.7B ternarized by ScaleQ-1.58 reaches over 90.52% of the performance of the prior best BitNet b1.58 2B4T averaged over 4 mathematics and coding tasks, and our ternary Qwen3-4B shows an absolute gain of 8.97%, while requiring 1,000,000 fewer calibration tokens for ternarization; (2) ScaleQ-1.58 generalizes well to 8 dense and MoE architectures, with performance significantly improving as model scale increases from 1.7B to 235B parameters; (3) ScaleQ-1.58 exhibits strong generalization across 13 tasks of varying difficulty levels, including challenging mathematics, coding and scientific logic reasoning, as well as basic commonsense reasoning and language generation. Intriguingly, ScaleQ-1.58 shows continuously improving performance as the number of calibration tokens increases. Besides, AYOT generalizes well beyond 1.58-bit quantization to other bit-width settings. With the default 4M calibration tokens, ScaleQ-1.58 requires only 4 to 240 hours on 8 A100-80GB GPUs to ternarize 1.7B to 235B reasoning LLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.