Can LLMs take the plunge? First, map the stairs
Abstract
Can large language models generate genuinely new hypotheses that may lead to scientific discoveries, or just combine the training data in a sophisticated way? We make this broad question testable by decomposing abduction (generation of explanatory hypotheses from observations) into a staircase of diagnostic steps. One path tests whether LLMs can detect and act on their own errors; the other climbs from selecting among supplied explanations, to recognizing structural correspondences, to constructing unsupplied hypotheses with testable consequences. For routine abduction, the tested LLMs underweight base rates relative to sample evidence, a pattern Grether (1980) documented in humans; across increasing model sizes, both fitted weights rise but their imbalance persists. Internal error signals become most informative where outputs are confidently wrong. In a toy model tracked across training, the internal signal becomes more informative than output uncertainty after a grokking-like transition. Across two frontier-model families, structural-recognition accuracy remains near chance in a single pass (0.50-0.54) but rises to 0.98-1.00 when hidden serial computation is permitted. In radical abduction without supplied candidates, when contradictory evidence signals a failure, models reach the correct hypothesis class in one world family, and occasionally produce correct final programs, but do not complete the full evidence-bound chain. Without such a trigger, two frontier-model families do not distinguish rule-bearing worlds from matched rule-free ones unless the hypothesis family is supplied. When we supply the method, a second pass repairs 110 of 114 first-pass errors. An interior reader can allocate retries but without a conclusive advantage over a trained output reader. Computation and supplied structure improve performance substantially, while constructing and testing explanations without those supports remains unreliable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.