acceptodds
Under review as a conference paper at ICLR 2027

Reason the Scene, Reconstruct the Space, Render the Unseen: TinkerCube for Spatial Intelligence

Abstract

Spatial reasoning, 3D reconstruction, and novel-view synthesis describe complementary aspects of the physical world, yet are typically isolated, limiting exchange between semantic and geometric evidence. We present TinkerCube, a coordinated multimodal model connecting spatial understanding, target-view geometry prediction, and controllable visual generation. TinkerCube augments a Mixture-of-Transformer-Experts backbone with a multi-view geometry pathway and role-specific interfaces that align 2D appearance, camera states, and cross-view structure in a shared context. Geometry-aware reasoning supplies explicit spatial evidence to the language expert; language-conditioned geometry converts motion and content instructions into target cameras, depths, and point maps; render-guided generation projects a reference-only point cloud into the target view to constrain completion. Three-stage training progressively aligns spatial tokens with language, learns motion-conditioned target geometry, and adapts coverage-aware view synthesis while preserving earlier capabilities. Across spatial question answering, depth and camera estimation, point-map reconstruction, and novel-view synthesis, TinkerCube delivers consistent performance, with leading results on major spatial-reasoning benchmarks and strong view-synthesis performance on DL3DV and RealEstate10K. These results show that integrating semantic reasoning, 3D geometry, and visual generation through joint training and interaction effectively strengthens spatial intelligence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.