acceptodds
Under review as a conference paper at ICLR 2027

Learning to Think and Explore in 3D World

Abstract

We study how vision-language models (VLMs) can learn to think and actively explore in 3D space. Instead of assuming informative input and providing multi-view images, videos, or complete 3D representations directly to the model, we preserve the native visual encoding of VLMs: a single image at a time, just as humans perceive the world. We treat the 3D representation only as an environment with which the model can interact through actions, such as moving left, and observations in the form of rendered images. Each rollout begins from a fixed initial first-person view together with a visual question. The model can then iteratively reason, navigate, and reorient its viewpoint, with each action inducing a new observation, before committing to a final answer. Following an initial stage of trajectory-supervised fine-tuning, we optimize the interactive policy with reinforcement learning (RL), using rewards driven primarily by final-answer correctness. Crucially, RL provides no direct supervision over which viewpoints the model should visit or which exploration actions it should take. Instead, the model must learn an effective exploration policy solely from the extent to which its interactions improve performance. Across six 3D benchmarks, our policy achieves the highest average score among active models at 54.1%, improving upon its base Qwen model by 22.9% and surpassing the proprietary Gemini 3.6 Flash model by 2.6% under the same interaction protocol. The evaluation includes four out-of-domain benchmarks, on which our policy achieves an average improvement of 7.2%. It also outperforms every evaluated open-source 3D model despite using 50–70% fewer input frames and lower-resolution renders. Further analysis shows that RL improves accuracy across all interaction levels by 3.2–6.2% and induces question-conditioned exploration: the policy searches for occluded targets, briefly confirms visible ones, and gathers multiple viewpoints for relational comparisons. Our RL design enables VLMs not only to reason over observed evidence, but also to decide when and how to acquire the evidence needed for generalizable 3D question answering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.