acceptodds
Under review as a conference paper at ICLR 2027

BenchMIRT: Disentangling Safety and General Capabilities in LLM Evaluation

Abstract

LLM benchmarks are designed to evaluate models along specific constructs; however, they often have confounders (e.g., item complexity, formatting requirements) that muddle the abilities that they actually measure. In this work, we leverage psychometrics theories to determine the latent factors that LLM benchmarks measure, focusing specifically on safety and general reasoning abilities, which are often seen as in tension. Applying Multidimensional Item Response Theory (MIRT) to LLM benchmarks, we introduce \modelname, a two-dimensional MIRT model fit on 100 LLMs' outputs across sixteen safety and general reasoning benchmarks. We find that our two-dimensional model consistently uncovers distinct safety and reasoning latent factors without any conditioning, and that these factors can facilitate accurate predictions of LLM scores. We further demonstrate that the latent safety and general reasoning item parameters (i.e., difficulty and discrimination) can be used to audit benchmarks, surfacing different ability requirements between subsets, isolating possible item errors, and showing that some benchmarks are dependent on a different ability than they are known to measure (e.g., BBQ). Our work showcases the utility of psychometrics models to audit benchmarks towards more construct validity in LLM evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.