acceptodds
Under review as a conference paper at ICLR 2027

UltrafastBench: A Deterministic Benchmark for Predictive Experimental Reasoning

Abstract

We introduce UltrafastBench to test final-answer prediction, diagnosis, and constrained control in specified ultrafast-optics systems. Its 400 questions span 24 physical-model families, including 240 design tasks. Executable references and typed answer fields support deterministic scoring without an LLM judge. On 380 unchanged screened prompts, five models answer 88–284 completely. On a frozen 400-question fresh-parameter cohort, GPT-6 Luna and GPT-6 Sol answer 115 and 303 without tools; isolated Python access raises these scores to 339 and 399, rescuing 233/285 and 97/97 prior misses. In four category-plus-estimate families, Sol chooses all 80 categories but completes only 49 answers closed-book, rising to 80 with Python. Saved traces show both numerical rescues and tool-induced errors. The results establish sensitivity to this Python-enabled protocol on new parameter worlds within the developed families. Versioned questions, references, and traces permit field-level replay.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.