acceptodds
Under review as a conference paper at ICLR 2027

SparseDiag: State-Dependent Diagnosis and Repair for Low-Bit Large Language Models

Abstract

Post-training quantization (PTQ) methods for large language models (LLMs) have largely advanced by changing how quantized layers are optimized, through improved objectives, learnable parameterizations, reconstruction granularities, and layer- or window-wise schedules. However, existing methods generally follow optimization schedules that are specified independently of the residual state that emerges as quantization progresses, leaving a complementary question underexplored: once a model reaches a particular quantized state, is every scheduled layer still worth optimizing? Through controlled layer-wise probing, we find that the marginal utility of further optimization is highly non-uniform and strongly state-dependent. In relatively weak quantized states, repairable errors can be broadly distributed across the network. After stronger PTQ optimization, however, many layers provide little additional benefit, while the remaining headroom often concentrates in a small set of non-contiguous residual bottlenecks. Moreover, repairing one bottleneck reshapes the subsequent repair landscape, suggesting that the layers worth optimizing cannot, in general, be determined reliably from the initial state alone. Motivated by these observations, we propose Sparse Diagnosis (SparseDiag), an iterative framework that explicitly separates diagnosis from correction. At each round, SparseDiag constructs lightweight layer-wise repair candidates using calibration data, ranks them on a disjoint calibration split, verifies the selected candidate on held-out calibration samples, and updates only a validated residual bottleneck. It then re-diagnoses the updated model and automatically stops when no reproducible repair opportunity remains. Consequently, the optimization schedule is discovered from the evolving quantized state rather than prescribed in advance. Experiments on Llama and Qwen models under W4A16 and W2A16 show that SparseDiag consistently exposes additional repairable headroom across different quantized starting states, while stronger PTQ states exhibit substantially sparser residual repair landscapes than weaker ones. Repeated diagnosis further reveals an evolving progression from broadly distributed repairability to sparse residual bottlenecks and eventually to headroom exhaustion. These results recast low-bit PTQ refinement from merely asking how to optimize each layer to deciding whether, where, and when further optimization is still beneficial.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.