What Do Benchmarks Measure? Discovering Benchmark Demands and Grounding System Capabilities
Abstract
Modern LLM benchmarks have grown increasingly complex and heterogeneous, yet what they actually measure remains unclear. Individual benchmark items often require multiple overlapping demands, while evaluation collapses performance across them into a single aggregate score. Existing fine-grained evaluation approaches typically specify capability dimensions in advance, but for complex benchmarks the relevant structure may itself be unknown. We therefore ask whether benchmark demands can be discovered from item semantics and grounded in system behavior, rather than specified in advance. To address this question, we introduce BenchScope, which separates demand discovery from behavioral grounding. BenchScope first uses a sparse autoencoder to discover recurring demand factors from benchmark items alone, without access to system responses. It then anchors an item response model to these fixed factors and uses system responses to test their behavioral relevance. The model separates overall system ability, semantic item difficulty, and demand-specific system specialization. This decomposition defines a shared, interpretable space for profiling benchmark demands and system capabilities. Across seven complex benchmarks, we recover reproducible and interpretable demand factors, with system performance varying systematically along these dimensions beyond overall ability and item difficulty. The resulting profiles reveal that similar benchmarks can emphasize different demands, while similarly scoring systems can exhibit distinct capability specializations. Together, these findings show that benchmark demand structure can be discovered from item content and grounded in system behavior, without requiring a predefined capability taxonomy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.