acceptodds
Under review as a conference paper at ICLR 2027

TernRefine: Gradient-Guided Fixed-Capacity Refinement of Ternary LLMs

Abstract

Post-training ternarization produces a compact discrete model, but the resulting ternary assignment need not be optimal for the target task. We introduce TernRefine, a refinement method that reallocates side-state occupancy after quantization while preserving the ternary codebook and the exact number of side states in every quantization group. TernRefine computes a single task gradient at the deployed quantized model, uses the first-order score −⟨G, ∆Q⟩ to rank legal relocations, and selects how many relocations to apply using a separate validation split. We find that the point at which the gradient is evaluated matters substantially. In a controlled study over five fitting offsets, gradients computed at the quantized model and at the corresponding full-precision model produce patches with less than 2% Jaccard overlap and lead to opposite average outcomes on held-out data. Ranking with the quantized model gradient improves WikiText-2 in all five runs, whereas ranking with the full-precision gradient degrades it on average. TernRefine also improves WikiText-2 and C4 on OPT-350M, Llama-2-7B, and Qwen3-8B after affine ternarization. Starting from PT2 models, all three fitted patches improve both datasets on Llama-2-7B and Qwen3-8B, and the selected models show positive average transfer on standard zero-shot tasks. Overall, these results show that ternary models can retain useful task structure after quantization and that gradients evaluated at the deployed discrete state provide an effective signal for refining that structure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.