Compress Locally, Predict Globally: Composable Margin Statistics for Estimating Compressed LLM Accuracy
Abstract
We introduce a statistical model that predicts the task accuracy of compressed language models without constructing and evaluating each candidate configuration. For each block and compression option, we summarize changes in answer margins (the score differences between competing answers) by a scaling coefficient and a residual variance. We estimate these statistics by compressing one block at a time, then reuse and combine them across configurations. We use the combined statistics to parameterize a Gaussian score model. For each example in a labeled reference set, we use the uncompressed model's scores to estimate the probability that the compressed model selects the correct answer. We average these probabilities to predict task accuracy. Across exhaustively evaluated MMLU quantization and pruning allocation spaces, prediction mean absolute error (MAE) ranges from 0.61 to 3.83 percentage points (pp). Across five K2-Horizon-7B quantization spaces, composing margin statistics achieves a mean per-space prediction MAE of 1.10 pp, compared with 1.55 and 1.42 pp for additive and multiplicative accuracy composition. Comparisons with covariance-based and empirical-residual predictors further support the accuracy estimates from the fitted Gaussian model. GPTQ prediction MAE remains within 0.67-1.19 pp across five checkpoints from 7B to 65B and reaches 0.90 pp on 24 presampled individual-layer allocations. The same prediction rule extends to BoolQ, HellaSwag, WinoGrande, and encoder-based AG News classification. As an application, maximizing predicted accuracy selects the best measured configuration in the complete Gemma-31B quantization space and improves pruning allocation over OWL under matched budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.