acceptodds
Under review as a conference paper at ICLR 2027

Beyond Accuracy: Re-evaluating Failures in LLM Physics Reasoning

Abstract

Evaluating physics reasoning in large language models (LLMs) requires more than identifying where models fail: it requires determining whether those failures are evidence of a recurring reasoning weakness. An observed error may disappear under repetition, a persistent error may not survive controlled reformulation, and a new failure may arise from a different reasoning breakdown. This distinction is especially consequential in physics, where errors in physical modeling, formula applicability, and symbolic or numerical execution can produce similar incorrect outcomes. We introduce **PACE** (**P**robe-based **A**ssessment and **C**haracterization of **E**rrors), a physics-grounded diagnostic framework that distinguishes *observed*, *stable*, and *attributable* failure evidence. **PACE** combines selective repeated measurement, mechanism-abstracted Difficulty Cards, controlled probe synthesis, scientific validation, and source-relative attribution to test not only whether a model fails again, but whether the same reasoning breakdown recurs. Across frontier reasoning models and multiple physics subfields, we find that repeated failure, failure on a source-guided probe, and broad error-category agreement each provide useful but incomplete diagnostic evidence. In particular, source-guided probe failures do not necessarily reproduce the source reasoning breakdown, and category-level similarity can conceal differences in the underlying reasoning breakdown. These results expose a gap between reproducing an error and diagnosing a recurring weakness, motivating a reliability-aware view of failure evaluation, in which evidence strength is measured rather than assumed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.