Evaluation Collapse: When Metrics Cannot Distinguish Model Behavior
Abstract
Scalar evaluation metrics (accuracy, AUROC, F1, perplexity) rest on three structural restrictions: pointwise scoring, permutation invariance, scalar aggregation. We prove this combination cannot represent behavioral properties, including detection delay, subgroup performance, and confidence trajectories, that depend on discarded information. For any such property, behaviorally distinct models receive identical metric scores, even at the maximum score where practitioners are most confident. No metric can recover what its representation discards. We prove: (1) collapse occurs whenever behavioral information is absent; (2) collapse families persist at optimal scores, making metric tuning futile; (3) for ordered sequences, collapse size follows an exact combinatorial law (Theorem 3). We instantiate this in three settings: anomaly detectors at differing by 50 steps in latency, classifiers with identical accuracy differing by 20 points in subgroup error rates, and language models with identical perplexity spanning full confidence-trajectory ranges. The phenomenon appears in real systems: COMPAS, a deployed recidivism-risk tool, exhibits this signature, with near-identical accuracy across race concealing a 20-point divergence in false-positive/negative rates, independently reproduced from public data. We operationalize the result as a lightweight audit on already-computed scores (no retraining), with characterized failure modes: it reliably detects collapse but can underestimate by if the perturbation distribution is too narrow. The core question: before certifying behavior from a metric, does the evaluation retain enough information to see what you're certifying?
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.