Before the First Equation: Selective Scene Control in Large Language Model Physics Reasoning
Abstract
We study how a solver's description of a physics scene shapes its answer. Across 284 questions from PHYBench, SciBench, and UGPhysics, four fixed model interfaces, and three repeated calls per condition, a question-matched six-field contract reaches 94.22% judged accuracy. This is +8.27 points above a no-contract scene baseline collected in a separate run (85.94%; descriptive). Within a shared response run, the matched contract exceeds a mismatched contract from another question by 74.15 points (Holm-adjusted paired sign-flip ). Its benefit depends on correspondence: a one-field solution-relevant edit in the contract yields 56.19%, compared with 79.28% for an edit designated irrelevant to the registered solution in the contract, a 23.09-point paired gap (Holm-adjusted ). The sensitivity is selective but not exclusive; even the designated irrelevant edit has a cost. In a same-call, solving with designed general six-direction self-check reaches 93.57%, versus 85.71% for direct solving (+7.86 points across separate runs, descriptive). Judgments use condition-specific mixtures of programmatic checks, model-assisted judging, and human QC. A well-matched physical description, and potentially an inline check of one's own solve, can support more accurate physics answers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.