Smaller Than the Noise: A Measurement-Theoretic Audit of Evaluation Protocols for LLMs
Abstract
Option-order sensitivity complicates multiple-choice evaluation of large language models (LLMs), but its magnitude must be interpreted relative to repeated-generation variability and measurement precision. We evaluate six open-weight LLMs on 25 text-based items from the International Cognitive Ability Resource under eight cyclic option rotations, with two greedy and five stochastic administrations per rotation, yielding 8,400 item responses. We estimate domain-specific abilities using a between-item multidimensional item response theory model with fixed parameters reconstructed from human response summaries and standardized factor loadings. A mixed-effects analysis separates shared rotation effects, model-specific rotation effects, and residual variation across draws. Under stochastic decoding (), the estimated order component has an SD of 0.08–0.10 ability units, compared with 0.30–0.33 across draws. Under greedy decoding, 99.7% of scoreable repeated item-response pairs are identical, yet the order component has an SD of 0.28–0.30. Thus, repeatability under a fixed presentation does not imply robustness to option order. Across both regimes, these pooled components remain below half the mean information-based conditional standard error of measurement. These comparisons are conditional on the reconstructed response model and do not establish equivalence between models. Our findings show how reporting order sensitivity alongside repeated-generation variability and test information can clarify what short evaluations support.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.