acceptodds
Under review as a conference paper at ICLR 2027

B-L-V: Where Reasoning Breaks Down in Large Language Models

Abstract

Large language models (LLMs) achieve strong performance on reasoning benchmarks, yet the sources of their failures remain opaque: accuracy and process reward models can tell us whether an answer is correct, but not where or why reasoning goes wrong. We observe that human reasoning relies on three functionally separable operations—the retrieval of task-relevant knowledge, the rule-based transformation of that knowledge into a conclusion, and the metacognitive verification of the resulting output—and that cognitive science shows these operations fail independently and in systematic, predictable ways. B-L-V turns this tripartite structure into a diagnostic lens for machine reasoning. It decomposes the standard answer of every task into Belief (the knowledge a solution presupposes), Logic (the derivation steps that yield a conclusion), and Value (the output requirements and verification a correct answer must satisfy), yielding a fixed, model-independent reference. Any model's free-form answer is then read against this reference, and each failure is attributed to the primary component where reasoning broke down. Across mathematics, commonsense, and value identification, B-L-V reveals that reasoning failures are not homogeneous: some stem from inaccessible knowledge, some from broken derivation, and some from neglected verification, with different models exhibiting different dominant failure loci—patterns that outcome-based metrics structurally cannot expose. B-L-V thus provides component-level, actionable, and reproducible diagnosis of where LLM reasoning breaks down.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.