Characterizing Abductive Reasoning Failures in Large Language Models with Elenchos
Abstract
Large language models (LLMs) are highly capable of reasoning from known rules to their consequences, but their ability to reason in the opposite direction – from observed behavior to the rules that produced it – remains poorly understood. Recent work has characterized this transition from observations to hypotheses about underlying rules as an “abductive jump” and argued that it remains challenging for current reasoning systems. To investigate this challenge, we introduce Elenchos, a framework for evaluating abductive inference in controlled executable formal environments, instantiated here in a dependently typed -calculus kernel. Given a pair of kernels, exactly one of which is sound and the other intentionally mutated, agents interact with both through a black-box interface to determine which of the two kernels is sound and identify the mutations responsible for observed behavioral differences. Evaluating frontier and mid-tier LLMs on Elenchos reveals a striking dissociation between detecting an alteration and identifying its cause: models frequently determine which kernel is sound and gather probes that uniquely distinguish competing mutations, yet often fail to identify the mutations responsible for the observed behavioral differences. To isolate the mechanisms driving this breakdown, we conduct diagnostic ablation studies on GPT-5.4 (high reasoning). These show that forward simulation of mutated systems is substantially harder than simulation of the sound system, but that simulation accuracy does not predict attribution accuracy during interactive reasoning. Remarkably, when the predicted consequences of competing hypotheses are provided explicitly, the model identifies the responsible mutation with near-perfect accuracy. Together, these results point to two interacting bottlenecks in abductive reasoning: deriving the behavioral consequences of competing hypotheses and integrating those consequences with accumulated evidence during interactive attribution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.