OpticalBench: A Physics-Grounded Benchmark for Vision-Language Reasoning on Optical Experiments
Abstract
Artificial intelligence is moving into the experimental loop of the physical sciences, proposing experiments, driving instruments and reading what they return. However, an optical system is still set up and aligned by hand, and whether a measurement is sound is judged by a human experimenter rather than by anything automatic. Vision-language models (VLMs) are a promising tool for such systems, given their increasingly strong reasoning. Yet no benchmark measures their capabilities on optical experiments across different levels of reasoning and various optical tasks. To bridge this gap, we introduce OpticalBench, a benchmark of setup–result pairs spanning six representative optical tasks, with each instance comprising the system components, their poses and parameters, the setup diagram, and the corresponding optical results. OpticalBench contains 1,961 scenes and 31,887 image inputs in total. Based on these setup–result pairs, we synthesise 7,664 question–answer pairs at four progressive levels of reasoning: Perception, Derivation, Prediction, and Decision, all grounded in physically meaningful quantities and relationships. We evaluate ten proprietary and open-weight deployments under a unified protocol, with per-item chance floors and controlled visual ablations. Experimental results reveal substantial differences across models: overall accuracy ranges from 13.3% to 65.0%, averaging 38.0% across the ten deployments. Performance also varies markedly across reasoning levels, averaging 47.9% for Perception, 39.3% for Derivation, 28.6% for Prediction, and 36.3% for Decision, revealing substantial room for improvement in reasoning over optical experiments. OpticalBench serves as a physics-grounded benchmark for measuring and advancing the capabilities of VLMs toward automated optical experimentation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.