acceptodds
Under review as a conference paper at ICLR 2027

Diagnostic Fine-Grained Ability Profiling for Large Language Model Evaluation

Abstract

Evaluating large language models (LLMs) is a prerequisite for characterizing their ability boundaries and for guiding subsequent model improvement and real-world deployment. Current practices, however, rely predominantly on aggregate benchmark scores, which conceal the heterogeneity of underlying abilities and exhibit inconsistent rankings across evaluation settings. To address this limitation, we propose a diagnostic framework that estimates multidimensional ability profiles directly from test item level responses. This framework dynamically constructs a structured Q-matrix that explicitly represents the associations between test items and the fine-grained abilities they require. It then integrates this mapping with multidimensional Item Response Theory (mIRT) to jointly estimate an interpretable ability profile for each model, rather than collapsing distinct abilities into a single aggregate score. Instantiated with a 35-dimensional mathematical ability taxonomy, this approach demonstrates strong criterion validity (ability scores aligning with observed accuracy on associated items), high cross-benchmark reliability (consistent model rankings across benchmark pairs), and accurate prediction on held-out items. Furthermore, it generalizes to evaluation in physics, chemistry, and computer science, preserving high validity and reliability under varying item distributions. Overall, our framework replaces the single aggregate score with a multidimensional ability profile, providing a principled basis for more reliable, ability-aware evaluation in support of model improvement and deployment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.