FluidBench: Diagnosing the Knowing–Doing Gap in Scientific AI through Fluid Mechanics
Abstract
As large language models and AI agents move from answering scientific questions toward solving real scientific problems, a critical question is whether scientific knowledge can be reliably translated into scientific action. Fluid mechanics offers a revealing testbed: success requires both applying physical knowledge to concrete problems and sustaining reliable computational research workflows. Existing benchmarks, however, typically evaluate scientific knowledge, problem solving, or research execution in isolation. We introduce **FluidBench**, a multi-level benchmark for diagnosing the *knowing–doing gap* in fluid mechanics. **FluidBench-Fundamental** contains 625 textbook-grounded problems assessing foundational knowledge; **FluidBench-Application** comprises 329 multi-stage problems with matched prerequisite probes and six-dimensional failure diagnosis; and **FluidBench-Research** contains 10 open-ended tasks across five research topics for evaluating end-to-end scientific agents. Experiments reveal two distinct gaps. First, a *knowledge-to-application gap*: models frequently answer all prerequisite probes correctly yet fail the original problem, with physical modeling emerging as the major bottleneck. Second, a *knowledge-to-research gap*: despite strong foundational fluid-mechanics performance, current agent systems remain far below reference-study quality, with failures concentrated in sustaining controlled experiments and complete, verifiable evidence. Together, these results show that strong scientific knowledge alone is insufficient for either reliable problem-specific reasoning or autonomous scientific research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.