The Models Behind the Claim: Empirical Standards for LLM Evaluation
Abstract
Selecting an evaluation panel is a consequential but largely unguided part of Large Language Model (LLM) research: we characterize the primary approaches researchers can use to determine how many and which models to include in an LLM evaluation, and we examine the advantages and limitations of each approach, including global standards derived from freshly published, year-long research on arXiv and local standards within the ICLR community. To see how these choices are made in practice, we trace model use across 11,428 arXiv papers that report LLM evaluations and compare this broader record with 486 eligible papers from the ICLR'26. Across this record, a consistent pattern emerges: most evaluations rely on a small set of models drawn from even fewer model families, so evaluation results may only represent a narrow portion of the LLM landscape; the median paper tests 3 models from 2 families on one task family, and 39.1% test a only single family. The ICLR comparison bolsters this point: papers associated with the venue tend to include more model variants, with most of the additional breadth coming from releases within the same families. Collectively, no single minimum model count can determine whether a panel is adequate; adequacy depends on what a paper claims, what it evaluates, how current its evidence is, and which model populations it intends to represent. We therefore translate these patterns into condition-specific standards for panel breadth, family and task coverage, recency, and model access. The resulting framework gives researchers a practical basis for planning LLM evaluation panels and gives the community a clearer basis for judging whether the evidence supports the claims.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.