Search or Experiment? Budget-Aware Metareasoning for LLM-Driven Scientific Discovery
Abstract
Large language models (LLMs) are increasingly used to automate scientific discovery. Yet, existing approaches typically optimize hypothesis search over fixed evidence or active experimentation over newly acquired evidence in isolation. These operations have sharply asymmetric costs, since hypothesis refinement consumes computation, whereas experiments require substantial resources. Moreover, the marginal value of each action changes throughout discovery and must be inferred from the evolving discovery state. The central challenge is therefore to continually determine whether the next unit of budget is better spent refining what is already known or acquiring evidence that can change what is known. We introduce CoREAgent (Cost-optimized Refinement and Experimentation Agent), a budget-aware framework that casts scientific discovery as sequential allocation between SEARCH and EXPERIMENT. At each step, a learned metareasoner estimates the cost-adjusted downstream value of each action from the discovery state and chooses between refining hypotheses and acquiring evidence, while a separate acquisition policy decides where to experiment. CoREAgent learns these values online through prequential credit assignment, scoring each decision on evidence acquired after it was made. Across four interactive discovery benchmarks, CoREAgent improves exact accuracy by 8 − 19 percent over the strongest baseline while reducing experimental usage by up to 31%. Our analyses isolate the contribution of each component and show that efficient discovery requires optimizing not only what to search for or experiment on, but when each action is worth its cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.