acceptodds
Under review as a conference paper at ICLR 2027

Conformal Risk Control for Sparsity Selection in Large Language Models

Abstract

Model pruning can effectively reduce the deployment cost of large language models (LLMs), while selecting an appropriate sparsity level is critical to balancing compression gains against performance preservation. Existing methods typically assess a prespecified sparsity level on the basis of its empirical performance on held-out evaluation data. However, this strategy lacks statistical reliability and may lead to uncontrolled performance degradation in real-world deployment. Therefore, we extend split conformal prediction, which can provide statistically valid guarantees for unseen samples under the exchangeability assumption, to pruning sparsity selection. Meanwhile, we propose CRCSS, a performance-degradation risk control framework for sparsity selection in LLM pruning. Specifically, we transform non-monotonic performance degradation trajectories into nondecreasing degradation prefix to determine sample-specific maximum safe sparsity boundaries, thereby formulating pruning sparsity selection as a conformal calibration problem. Extensive experiments across multiple LLMs, pruning methods and language understanding benchmarks demonstrate that CRCSS achieves empirical guarantee rates consistent with the prespecified reliability targets, while enabling a controllable trade-off between reliability and compression.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.