SPhyR: Environment for Spatial-Physical Reasoning with Simulation Verifier
Abstract
Reinforcement learning from verifiable rewards needs verifiers, and the available ones are narrow: exact match, unit tests, or a learned reward model. Each assumes a task has a checkable canonical answer. Many design problems do not. We present SPhyR, an environment for spatial-physical reasoning whose verifier is a physics simulation rather than a comparison to a stored answer. An episode presents a 2D domain with prescribed loads and supports and a masked material distribution; the agent submits a completion, and the environment solves it as a linear elastic structure and scores it against a topology optimizer re-run on exactly the masked region. A completion that routes the load differently from the reference but just as efficiently scores just as well. Across the valid completions in our results, carry their load correctly while differing from the reference somewhere; a string-matching verifier rejects them all, and graph connectivity disagrees with compliance on of them, so a cheaper structural proxy will not substitute. We show the two views rank models differently and in both directions, that the environment's reward is not reachable by retrieval, and that a benchmark of this shape can rank models by grid-serialisation ability rather than by reasoning unless that is measured separately. The environment is released in the OpenEnv interface, single-step or with simulator feedback between attempts (Footnote: The dataset, environment and evaluation code are released; links are withheld for double-blind review and the code accompanies the submission as supplementary material.).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.