MDGym: Benchmarking AI Agents on Molecular Simulations
Abstract
The promise of AI-driven scientific discovery hinges on whether AI agents can autonomously design and execute the computational workflows that underpin modern science. Molecular dynamics (MD) simulation provides a demanding test of this capability: it requires translating a physical problem into syntactically and semantically correct input scripts, choosing force fields, ensembles, and initial and boundary conditions, diagnosing engine errors and unstable trajectories, and interpreting outputs against known physical behavior. We introduce , a benchmark of 200 expert-curated MD simulation problems in and , two widely used MD engines, organized into three difficulty levels and comprising 366 target quantities scored against reference values from expert simulations. We evaluate three general-purpose command-line agent frameworks, Claude Code, Codex, and OpenHands, each with four LLMs. The best configuration solves 68% of easy problems, 30% of medium problems, and 21% of hard problems, and the choice of model changes performance far more than the choice of framework. Trajectory and engine-error analyses show that most engine errors arise while agents assemble the simulation input, that weaker models fail early through syntax errors and premature termination, and that stronger models complete more simulations but report incorrect values from error-free runs. Current agents therefore lack three capabilities required for autonomous simulation: coherent specification of simulation inputs, domain-aware recovery from engine errors, and physical and methodological validation of results.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.