acceptodds
Under review as a conference paper at ICLR 2027

Typicality of Representations Predicts Input-Level Performance in Language Models

Abstract

Predicting whether a model will perform well on a set of inputs remains a topic of central importance in machine learning. Motivated by the information-theoretic notion of typicality, we address this open problem by introducing representational typicality: a framework for measuring how well the intermediate representations induced by an input conform to a model's expected representational regularities, and consequently of the model's performance on the input. Without requiring any dataset labels or model retraining, the proposed method first derives the characteristic statistics of a model's intermediate representations using only readily accessible data. The method then evaluates whether the representations induced by new data are compatible, in the sense of being typical, with these characteristic representation statistics to predict input-level performance. Across an array of widely used language models and downstream tasks, our experiments comprehensively demonstrate that typicality is highly predictive of input-level performance, outperforms baseline measures, remains predictive across varying experimental configurations, and is robust to a suite of hypothetical confounding factors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.