SkillEval: Learning Interpretable Ability Profiles of LLMs via Cognitive Diagnosis Models
Abstract
Current evaluation of LLMs on benchmarks typically aggregates performance over individual examples into a single scalar score. This approach is overly coarse and ignores the diverse scope of skills covered by benchmark datasets. To address this limitation and better exploit benchmark data, we propose SkillEval, a new evaluation framework that performs cross-model, skill-level analysis. The framework integrates advances from two areas: (1) it leverages the meta-cognitive capabilities of LLMs to automatically discover interpretable skills from questions, and (2) it applies cognitive diagnosis models (CDMs) from psychometrics to estimate per-skill abilities of LLMs. We conduct experiments on 3,811 LLMs evaluated on 9,523 items across five benchmarks spanning diverse domains (MATH, BBH, GPQA, MuSR, IFEval). Compared to the traditional average-based evaluation paradigm, SkillEval produces reliable, more calibrated and informative yet concise model profiles that more accurately predict model behavior on unseen items. These profiles also reveal how benchmarks share skills and how instruction-tuned models differ from their base models. We also demonstrate the utility of SkillEval in two downstream tasks: (1) Computerized Adaptive Testing (CAT), where an unseen model can be efficiently evaluated using a trained CDM, and (2) LLM routing, which dynamically assigns each query to a capable but efficient model, reducing inference cost by over 60% while achieving improved performance. Code is provided as supplementary material to enable reviewer reproduction; the Q-matrix, skill taxonomy, trained NCDM parameters, and per-LLM mastery profiles will be released publicly upon publication.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.