A Distributional Distance Perspective on Evaluating LLMs
Abstract
Large language models (LLMs) are commonly evaluated by accuracy on public benchmarks. Since different LLMs are trained on different data distributions, raw accuracy may conflate model capability with training-data proximity, leading to unfair evaluations. The first factor is the well-studied data contamination. This paper explores another factor triggering unfair evaluation from the perspective of distributional distance between the input data and the LLM training data. We first adopt simple approximation methods to measure the distributional distance, and then empirically verify a strong statistical correlation between accuracy and distributional distance. We therefore introduce a two-dimensional criterion that jointly considers accuracy and distributional distance, enabling fairer evaluations. Finally, we re-evaluate nine LLMs on thirteen standard benchmarks and obtain several interesting findings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.