acceptodds
Under review as a conference paper at ICLR 2027

FrontierPhysics: Evaluating Agents for End-to-End Physics Research

Abstract

Large language models (LLMs) are increasingly participating in substantial parts of the physics research process. However, current physics benchmarks test fixed questions, isolated subtasks, or paper reproduction with given solutions, and rarely grade research quality alongside correctness. We introduce FrontierPhysics, a benchmark evaluating the ability of AI agents to conduct end-to-end research across theoretical, computational, and experimental physics and nine domains. The benchmark comprises 56 tasks made by 42 physics researchers, all from task authors' personally conducted research. In each task, agents are asked to write a scientific manuscript as well as structured output files. To succeed in these tasks, agents need to conduct in-depth research on the given topic to formulate proper research plans, select right methodologies, and produce a publishable manuscript of their results. For comprehensive evaluation, FrontierPhysics combines deterministic verifiers with rubric-based assessment. The former validates structured outputs against reference oracles, while the latter evaluates the agents' produced artifacts and trajectory according to both expert-authored rubrics evaluating the publish-ability of results and universal rubrics checking behavioral alignment. This evaluation protocol assesses the accuracy of the computation, scientific rigor, physics understanding, and scientific manuscript writing. Across eleven frontier agents, the best passes only 22% of rollouts, and seven pass at most 3%, indicating end to end physics research remain challenging to current models. Additionally, we conduct a detailed human study of the agents failure taxonomy by manually analyzing trajectories and identify failure modes. We find lack of rigorous thinking and expert intuition, instead of physics knowledge gap or execution mistakes as being the main weaknesses of frontier LLM agents on our benchmark.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.