ParticleBench: Evaluating LLM Agents on Experimental Particle Physics Tasks
Abstract
We introduce ParticleBench, a benchmark for evaluating large language model (LLM) agents on computational tasks in experimental particle physics. The benchmark focuses on long-horizon workflows that combine domain-specific reasoning, code development, and iterative numerical experimentation. Agents receive task specifications, data, and access to scientific computing tools, and must develop executable analysis procedures rather than provide only textual answers. Tasks are defined by computational objectives together with explicit requirements for the validity of their results. Evaluation separates correctness from performance: task-specific verifiers check whether submitted solutions satisfy prescribed numerical and physics-based requirements, while scoring measures analysis quality or computational efficiency among solutions that meet these requirements. Final submissions are frozen and executed on held-out data, with evaluation samples excluded from solution development and optimization. This protocol assesses whether improvements remain valid beyond the data available to the agent during development. ParticleBench provides an executable evaluation setting for studying how LLM agents implement and refine multi-step scientific analyses, with an emphasis on verifiable task completion under experimentally meaningful constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.