FlowIntentBench: Can LLM Agents Analyze Flow Fields to Answer Scientific Questions?
Abstract
Answering scientific questions about numerical flow fields requires choosing an appropriate analysis and determining what the resulting evidence supports. Evaluating LLM agents on this task is challenging because different valid analyses can yield different findings from the same data. We introduce FlowIntentBench, a benchmark of 24 scientific question families across 15 flow-field datasets. Each family is instantiated in four conditions that vary which analysis decisions the question specifies and how findings are to be selected and reported, and is paired with a set of accepted operationalizations (definitions, measurements, and decision rules) together with the findings each yields. Responses are scored on the validity of the stated analysis, on recall of the findings it requires, and on the consistency between findings and operationalization. Across nine tool-using LLM agents, every model scores lower when all principal analysis decisions are left to it. Mean validity falls from 0.846 to 0.742 and recall from 0.910 to 0.863. Our results show that current agents perform reliably when an analysis is prescribed, but struggle when they must determine how to analyze the scientific question.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.