acceptodds
Under review as a conference paper at ICLR 2027

Screening Is Not Ranking: Resolution-Aware Model Selection Under Compression

Abstract

Fidelity metrics such as reconstruction error, perplexity, and teacher agreement are effective at detecting severely damaged compressed models, but they need not identify the best model among candidates that all retain high fidelity. We therefore formulate compressed-model selection as a decision problem and evaluate metrics through selection regret: the capability gap between the selected model and the best available candidate. We identify four structural limits on whether a proxy can achieve low regret: a proxy limit, when fidelity bounds are too loose to order viable candidates; a pool limit, when compression rate is confounded with design quality; a target limit, when the evaluation budget is insufficient to resolve capability differences; and a target-definition limit, when no single scalar capability target is appropriate across deployments. We derive a formal result for each limit and test them on pre-registered pools of lexical-interface compressed language models. Fidelity is highly reliable for screening broken models (AUROC 0.97–1.00), yet in a viable 3B-model pool, bytes and nine fidelity metrics recover essentially the same ranking because the pool collapses to a compression-rate ladder rather than distinguishing design quality. Evaluation budget sharply determines what can be resolved: across 11,734 pair-by-budget tests, our regret bound is never violated, and pairs classified as resolvable reverse ordering only 0.005% of the time, compared with 22.8% for unresolved pairs. Screening before capability evaluation reduces selection regret by up to an order of magnitude. More importantly, we observe direct proxy reversals at fixed compression rate: restoring row norms improves capability by +0.0089 across eight paired seeds while every measured fidelity proxy ranks the modified model worse; a pre-registered SmolLM2 replication yields a +0.0050 capability gain in 11/12 seeds while 4/5 fidelity proxies again prefer the baseline. These results support a simple evaluation protocol: use fidelity for screening, control for compression rate and model family, evaluate capability among the survivors, and report “unresolved at this budget” when the available evidence cannot support a ranking.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.