acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Item Response Theory for LLM Evaluation with Item-Marginalized Maximum Likelihood

Abstract

Item Response Theory (IRT) is increasingly used to rank language models, analyze benchmarks, and reduce evaluation cost. Yet conventional IRT estimation was developed for a response matrix geometry with many examinees and comparatively few items. Modern LLM evaluation often reverses this regime, comparing a small number of models on hundreds or thousands of benchmark items. We study the consequences of this mismatch and develop item-marginalized IRT estimation as an alternative formulation which treats evaluated model abilities as fixed inferential targets and benchmark items as random effects. Across simulations calibrated to six benchmarks, we find a systematic geometry-dependent crossover. Conventional person-marginalized IRT performs better when models are abundant, whereas item marginalization becomes increasingly advantageous as the response matrix becomes item rich. With only 12 models, item marginalization reduces average Kendall ranking loss across six benchmarks by 56% relative to person marginalization. This advantage translates directly to efficient LLM evaluation. With only 12–40 calibration models, item-marginalized item selection reduces held-out ranking loss by 18%, enabling IRT-based benchmark distillation from only a few dozen fully evaluated models rather than requiring large historical response matrices. On real sparse response matrices, item marginalization also consistently improves ranking recovery over conventional IRT. Mechanism and robustness experiments attribute these gains to improved population generalization of item characteristics and show that the crossover persists across alternative IRT models and strongly non-Gaussian latent populations. These results suggest that the direction of marginalization should reflect the geometry and generalization target of the evaluation, making IRT better aligned with modern LLM evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.