acceptodds
Under review as a conference paper at ICLR 2027

Learning Multi-Step 3D Reasoning for Embodied Question Answering

Abstract

Embodied Question Answering (EQA) requires embodied agents to explore a 3D environment to find evidence and answer questions related to the environment. Some methods employ off-the-shelf vision-language models (VLMs) to solve EQA tasks, which exhibit limited 3D reasoning capabilities, making it difficult to accurately retrieve evidence and resulting in inefficient exploration and incorrect responses. In this paper, we propose a chain-of-exploration EQA agent that unifies exploration, spatial evidence acquisition, and question answering through iterative reasoning with tool use. Specifically, our agent reasons over both current observations and historical information, determines what spatial information is required for solving the task, selects appropriate tools to acquire geometric evidence, and decides where to explore next. To improve the reasoning and exploration behaviors of the agent, we further design an EQA data generation pipeline that automatically constructs EQA tasks with 3D reasoning trajectories. Based on the pipeline, we collect the EQA-RT dataset, which contains 18K tasks and is divided into a training set, EQA-RT-Train, and two test sets, EQA-RT-Seen (scenes overlapping with the training set) and EQA-RT-Unseen (unseen scenes). We perform supervised fine-tuning on EQA-RT-Train to initialize the model's tool-use capability, followed by evidence-grounded reinforcement learning to enhance task-relevant spatial evidence acquisition. Across public benchmarks, our method improves LLM-Match by 0.69 points and grounding-aware success by 2.93 points, while reducing exploration distance by up to 61.4%. On our EQA-RT, it improves target-object recall by 0.06–0.12 and LLM-Match by 4.19–5.63 points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.