Ternova: Adaptive Low-Rank Post-Training Ternarization for Large Language Models
Abstract
Large language models (LLMs) have achieved remarkable capabilities, yet their massive memory footprint and inference cost hinder efficient deployment. Ternary post-training quantization (PTQ) offers an attractive solution by combining compact storage with richer representation capacity than binary quantization. However, existing ternary PTQ methods remain challenged by limited ternary representation capacity, heterogeneous structural sensitivity across Transformer modules, and the mismatch between local reconstruction and model-level objectives. To address these challenges, we propose Ternova, an adaptive low-rank post-training ternarization framework for LLMs. Ternova is built on three core components: 1) Hessian-Aware Annealed Ternary Factorization (HATF) constructs high-fidelity low-rank ternary representations through sensitivity-aware factorization and annealed discrete refinement; 2) Sensitivity-Reweighted Mixed-Rank Allocation (SRMA) adaptively distributes rank capacity according to structural irreplaceability, task sensitivity, and storage cost; and 3) Task-Aligned Ternary Refinement (TATR) further aligns the compressed model with task-level behavior through global scale calibration and sparse topology refinement. Ternova preserves a compact ternary deployment form without additional dense high-precision branches. Extensive experiments across diverse LLM families demonstrate that Ternova achieves competitive or superior performance over existing low-bit PTQ methods at only 1.59 BPW while retaining competitive reasoning capabilities. On Qwen3.5-27B, Ternova achieves 74.62% average zero-shot accuracy, approaching the 75.97% BF16 baseline. Moreover, Ternova achieves roughly 3 inference speedup and up to 80.6% memory savings over BF16, demonstrating a strong balance among accuracy, compression, and practical efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.