acceptodds
Under review as a conference paper at ICLR 2027

What Makes Agents Effective for Scientific Design? A Controlled Evaluation with Scientific Foundation Model Tools

Abstract

Scientific agents have shown promise for automating complex computational workflows, but their usefulness for scientific design hinge on reliably completing a task and finding high-quality solutions. Here, we evaluate how agent design choices—including the agentic harness, Large Language Model (LLM), and tool-calling mechanism—affect these two outcomes. We introduce Nomad, a platform for serving Scientific Foundation Models (SciFMs) through either direct tool calling or Programmatic Tool Calling (PTC) via the Model Context Protocol, allowing us to vary the tool-calling mechanism without modifying the tools or agentic harnesses. Our evaluation includes 845 agent rollouts across two multi-objective design tasks: designing a multi-material system for Richtmyer–Meshkov Instability formation and screening molecular space for liquid-electrolyte candidates. We measure design quality using Pareto hypervolume and fit a Bayesian hurdle model to separately estimate the probability of viable designs and the performance of the proposed designs conditional on success. We find that PTC most clearly improves reliability and performance for the multi-material design task, while the effects of harness and LLM are generally smaller and depend on the task and their pairing. Providing historical designs increases conditional hypervolume by 106% for the multi-material task, with agent traces suggesting that these examples help keep search near the surrogate model’s training distribution. Finally, higher-fidelity multi-physics simulations confirm that selected agent-proposed designs increased the historical Pareto-front hypervolume by +45% despite imperfect surrogate calibration. Collectively, our results demonstrate the importance of tool-calling mechanisms and scientific context when connecting SciFMs to agents for design tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.