acceptodds
Under review as a conference paper at ICLR 2027

SimVerse:Benchmarking Inverse Mental Simulation via Interactive Games

Abstract

Inverse mental simulation, the process of reasoning backward from observed physical and spatial outcomes to their causal antecedents, requires compositional inference over continuous physical dynamics and discrete spatial topology. Existing benchmarks predominantly target forward prediction, and none is dedicated to the inverse, effect-to-cause direction in multimodal large language models. We introduce **SimVerse**, a benchmark for inverse mental simulation in multimodal large language models. **SimVerse** comprises 2,486 simulator-verified instances across four interactive game environments, organized into a taxonomy by operation order-sensitivity (order-invariant or order-sensitive) and state-transition type (discrete or continuous). Every model output is executed in a domain-specific physics or geometry simulator, so any valid antecedent earns credit and partial credit reflects physical or geometric progress.We evaluate seven models against an IQ-screened human baseline (). The strongest model (Gemini-3.1-Pro) achieves an overall score of 52.7%, falling 22.8 percentage points (95% CI: 19.3–26.4) below the human reference (75.5%), with marked variation across task families. On dice face inference, models identify up to 92.3% of symbols, yet none exceeds 32% once orientation is also required, exposing a bottleneck in cumulative state tracking. A concise direct-answer prompt and removal of the rendered image each lower success for all seven models and raise it on no task, while per-task model rankings remain largely stable. For both humans and models, order-sensitive discrete tasks are markedly harder than order-invariant continuous ones, identifying operation order-sensitivity as a central obstacle for inverse reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.