SciConvBench: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science
Abstract
Large Language Models (LLMs) are increasingly studied as scientific assistants, with benchmarks assessing knowledge, reasoning, code generation, and tool use. These evaluations, however, typically assume the scientific problem is already well-posed, whereas practical scientific assistance often begins with an ill-posed user request that must be refined through dialogue before any computation, analysis, or experiment can be carried out reliably. We introduce SciConvBench, a benchmark for multi-turn clarification in scientific task formulation across four computational science problem domains: fluid mechanics, solid mechanics, materials science, and partial differential equations (PDEs). It is designed to evaluate two complementary capabilities: eliciting missing information (disambiguation) and detecting and correcting erroneous requests containing internally contradictory information (inconsistency resolution). Our benchmark combines a detailed task ontology with a rigorous rubric-based evaluation framework to deliver reproducible results and quantify LLM performance across clarification behavior, conversational grounding, and final-specification fidelity. Current frontier models perform relatively well on inconsistency resolution, but even the best model resolves only 52.7% of the disambiguation cases through dialogue, averaged across the domains. We further find that frontier LLMs often make silent assumptions and perform implicit repairs that are not grounded in the conversation. SciConvBench establishes a foundation for evaluating the upstream conversational reasoning that a reliable computational science assistant requires.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.