LLM Evaluation is a Forecasting Problem
Abstract
Standard procedure in machine learning is to randomly split train and test data, emulating IID draws from one underlying data distribution. In practice, models are trained on past data and deployed in the future, which may come from distinct distributions. In this paper, we observe that text distributions shift continuously over time and thus, because the goal of model evaluation is to estimate how well a model will perform in the future, we argue that evaluation should reflect the temporal structure of the data. Indeed, we find that web text is highly non-stationary, as documents grow more compressible and semantics, domains, and embeddings shift over time. We further observe that these shifts are temporally local, and model performance changes continuously over time across diverse benchmarks; consequently, future model performance is better predicted by measurements made on recent evaluation data, and is best predicted by time-series forecasters applied to historic performance rather than by historic performance itself. Choosing which model to deploy based on forecasted accuracy enables more reliable selection of performant models, thus improving downstream accuracy in the regime practitioners care about: deployment. Taken together, our experiments suggest that time should be treated as a fundamental variable in LLM evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.