acceptodds
Under review as a conference paper at ICLR 2027

SPhyR: Environment for Spatial-Physical Reasoning with Simulation Verifier

Abstract

Reinforcement learning from verifiable rewards needs verifiers, and the available ones are narrow: exact match, unit tests, or a learned reward model. Each assumes a task has a checkable canonical answer. Many design problems do not. We present SPhyR, an environment for spatial-physical reasoning whose verifier is a physics simulation rather than a comparison to a stored answer. An episode presents a 2D domain with prescribed loads and supports and a masked material distribution; the agent submits a completion, and the environment solves it as a linear elastic structure and scores it against a topology optimizer re-run on exactly the masked region. A completion that routes the load differently from the reference but just as efficiently scores just as well. Across the valid completions in our results, carry their load correctly while differing from the reference somewhere; a string-matching verifier rejects them all, and graph connectivity disagrees with compliance on of them, so a cheaper structural proxy will not substitute. We show the two views rank models differently and in both directions, that the environment's reward is not reachable by retrieval, and that a benchmark of this shape can rank models by grid-serialisation ability rather than by reasoning unless that is measured separately. The environment is released in the OpenEnv interface, single-step or with simulator feedback between attempts (Footnote: The dataset, environment and evaluation code are released; links are withheld for double-blind review and the code accompanies the submission as supplementary material.).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.