AbductionBench: A Unified Framework for Evaluating Abductive Reasoning in Large Language Models
Abstract
Abductive reasoning is the inference of plausible explanations for observed evidence, particularly when information is incomplete or supports multiple competing hypotheses. Evaluating this capability requires more than a single task: a system may need to generate an explanation, discriminate among competing hypotheses, revise its hypotheses as evidence arrives, or decide what evidence to seek next. Yet these capabilities are typically evaluated through isolated benchmarks with different task formulations and scoring procedures, so claims that a model is “good at abduction” may reflect only a narrow aspect of this capability. We introduce AbductionBench, a unified framework that brings this fragmented landscape into a common evaluation space. Its taxonomy characterizes abductive tasks by hypothesis format, application domain, and evidence-delivery mode, covering generation and selection tasks across static, passive, and interactive settings while preserving their task-specific execution structure. We instantiate this framework on 27 datasets spanning ten domains and evaluate eight LLMs from four model families—GPT, Gemini, Gemma, and Qwen—alongside Jev, a non-LLM System One decision system evaluated on single-choice selection tasks. Beyond harmonizing heterogeneous tasks, AbductionBench separates final-answer evaluation from analysis of the observable reasoning traces and interaction histories produced alongside those answers. We analyze written traces as observable model outputs rather than faithful representations of internal computation, characterizing evidence use, hypothesis comparison, uncertainty, contradiction handling, and answer commitment. Chain-of-thought is not uniformly beneficial across models and tasks, and within datasets, longer traces are associated with lower accuracy. Jev matches or exceeds small LLMs when selecting among explanations over compact evidence but falls substantially behind on longer cases, suggesting that candidate discrimination and extended evidence integration place different demands on a system. In interactive tasks, the proportion of actions judged relevant declines at later steps, and a relevant first question is associated with a correct final answer; in the passive task, models rarely revise their initial hypothesis as new evidence arrives. Together, these findings highlight the value of evaluating abductive reasoning through both final answers and observable evidence use, hypothesis revision, and information gathering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.