From Flat Scores to Capability Profiles: A Capability-Cognition Framework for Generative AI Benchmarking
Abstract
Scalar benchmark scores obscure which capabilities a benchmark samples and at what cognitive depth. We introduce Capability-Cognition Profiling (CCP), which represents a benchmark as a distribution over domain content and cognitive operations , and characterises it through five profile descriptors, content breadth, cognitive breadth, cognitive depth, cognitive evenness, and profile overlap, plus two corpus-contribution measures (exclusive and Shapley coverage contribution). CCP does not rank benchmarks by quality; it makes explicit the capability distribution over which benchmark scores are computed. We instantiate CCP on world culture by constructing a multilingual consensus taxonomy (1,291 leaves, 15 depth levels; Chu–Liu–Edmonds arborescence over Wikipedia language editions) and mapping 29 major cultural benchmarks into the resulting capability space via a two-model LLM jury with an open-weight arbiter and independent human validation at multiple resolutions. Across seven contemporary models on seven benchmarks (9,800 evaluations), joint profile distance is significantly associated with pairwise rank disagreement under an exact benchmark-level permutation test (, exact ; leave-one-benchmark-out mean ), while content-only and cognition-only distances do not individually reach significance (); the joint's advantage over each marginal is itself not statistically established (). Composition standardisation to a common target produces one nominal leaderboard change, although uncertainty is substantial. The formalism and the graph-consensus taxonomy recipe are both domain-general; only the Wikipedia instantiation is culture-specific.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.