When Rules Are Hidden: A Paired Evaluation of Abductive Reasoning
Abstract
Large language models (LLMs) are increasingly used for scientific discovery, where they must not only reason from known rules but also infer hypotheses and latent mechanisms from observations. Abductive reasoning is therefore a central capability, yet its evaluation faces two challenges. First, abduction is confounded by deduction: an incorrect prediction may arise from inferring an inadequate hypothesis or from reasoning incorrectly from it. Second, finite observations can admit multiple plausible hypotheses, making evaluation against a single ground-truth explanation problematic. We introduce RuleCast, a paired benchmark that addresses both challenges. Each problem is posed under Complete-Rule and Hidden-Rule conditions that differ only in whether the latent rules required for prediction are provided or must be inferred from evidence, so their gap characterizes the additional difficulty of inferring missing premises. Problems are built on controlled transition-system worlds in three scientific domains, where multiple evidence-consistent hypotheses may exist but agree on every evaluated forecast, removing the need for a unique ground-truth hypothesis. Across open-weight and proprietary reasoning models, forecast accuracy consistently drops when rules must be inferred, even for models that execute the given rules nearly perfectly. An error analysis over the enumerable hypothesis space locates the bottleneck in verifying candidate hypotheses against the evidence rather than in executing a committed hypothesis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.