More Than Meets the Eye: Measuring Logical Reasoning Beyond a Single Accuracy
Abstract
An increasing number of large language models (LLMs) have achieved high accuracy on existing logical reasoning benchmarks, demonstrating impressive performance. However, producing the correct answer does not necessarily imply that a model can reliably identify and respond to the logical relations underlying a problem. Recent studies have highlighted the limitations of using overall accuracy as the primary measure of logical reasoning ability and have turned to examining models' reasoning behavior. In this work, we introduce ReliLogic, a behavioral reliability evaluation framework based on propositional-logic entailment questions. The framework consists of a Representation Set containing four logically equivalent forms and an Intervention Set constructed through logic-preserving and logic-altering interventions. Beyond final accuracy, ReliLogic evaluates cross-representation reliability, selective logical responsiveness, and instance-level prediction stability, providing a more comprehensive characterization of logical reasoning ability from the perspective of model behavior. We conduct a series of evaluations across 74 language-model configurations and uncover multiple failure modes that are not reflected by accuracy alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.