Artemis-3D: Latent Spatial Reasoning for Visual-Spatial Intelligence
Abstract
Visual-spatial intelligence poses significant challenges for current multimodal large language models (MLLMs), whose language-centric representations limit their ability to reason about spatial relationships. Existing approaches address this limitation by either injecting geometric information through external modules or verbalizing spatial information into textual Chain-of-Thought (CoT), yet both ultimately retain a language-oriented reasoning space. Meanwhile, we observe that language-oriented visual representations can conflate semantically similar objects at distinct spatial locations, obscuring the spatial distinctions essential for reasoning. To overcome the constraints of language-oriented spatial reasoning, we introduce Latent Spatial Reasoning, a new paradigm that enables MLLMs to construct and reason directly within a spatially grounded latent space. We instantiate this paradigm in Artemis-3D, a unified framework with two complementary components. First, Mental Map Imaging (MMI) internalizes geometric structure into visual token representations, constructing a latent mental map that preserves spatial distinctions. Second, Visual-Cue Latent Reasoning (VCLR) autoregressively retrieves question-relevant spatial evidence from this mental map and performs variable-length reasoning directly in the spatial latent space. Together, MMI and VCLR enable Artemis-3D to preserve geometric structure throughout the reasoning process without explicit textual CoT. Extensive experiments demonstrate that Artemis-3D outperforms state-of-the-art methods on visual-spatial intelligence benchmarks. Our code will be publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.