CFI: Capability Fisher Information for Benchmark Diagnosis in LLM Evaluation
Abstract
In current practice, LLM benchmarks are often evaluated by the reliability of the model rankings they produce, yet stable rankings do not necessarily imply sufficient and balanced measurement of the capabilities they are intended to assess. We propose Capability Fisher Information (CFI), a measurement-theoretic framework for diagnosing benchmark structure through the distribution of information in a benchmark-specific capability space. CFI represents each item as an information contribution characterized by a capability direction and strength, aggregates these contributions into a benchmark-level information operator, and supports three complementary diagnostics: Coverage, Uniformity, and Redundancy. Together, they expose capability blind spots, information imbalance, and repeated probing that overall accuracy or item counts cannot reveal. Across text-only and multimodal benchmarks, CFI reveals structural deficiencies overlooked by conventional statistics, while controlled structural perturbations and held-out model experiments validate the effectiveness of its diagnostics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.