IdeaPrism: Beyond Ground-Truth Matching for Open-Ended Hypothesis Evaluation
Abstract
AI scientists generate open-ended hypothesis sets in which several scientifically plausible explanations may fit the same evidence, including mechanisms absent from a predefined target hypothesis. Many existing evaluations, however, score outputs by matching them to designated discoveries or by asking human or LLM judges for direct quality assessments. Direct ratings introduce evaluator dependence, while target matching can miss alternative hypotheses and obscure the breadth of a generated set. We introduce IDEAPRISM, a framework for open-ended evaluation of scientific hypothesis generation. IDEAPRISM characterizes hypothesis sets along three dimensions. Exploration Novelty measures how rarely a mechanism appears in a fixed reference population, with scores adjusted for estimated reference coverage. Diversity measures the effective number of scientific directions represented in a hypothesis set. Value measures the change in a fixed reasoner's accuracy on hidden-evidence questions when conditioned on a hypothesis. We evaluate IDEAPRISM on 61 topics spanning 13 scientific subfields. Controlled experiments assess reference-coverage estimation, effective-direction recovery, and hypothesis-conditioned prediction. We then apply the framework to ten AI-scientist scaffolds. The resulting profiles show that improvements in reference-relative exploration or predictive utility do not necessarily coincide with broader exploration across scientific directions. IDEAPRISM complements target-recovery benchmarks by providing an operational, multidimensional characterization of generated hypothesis spaces.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.