RoboQuest: Generalist Physical Agents that Search, Inspect and Test
Abstract
Recent advances in multimodal foundation models have them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction should it be absent in the observations. It may need to determine where a relevant object is, inspect an unobserved property, discover the effect of an unfamiliar tool, or manipulate the environment to make the required information accessible. We thus introduce a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. The benchmark centers on three forms of uncertainty—search, manipulation-based inspection, and interactive testing—while allowing viewpoint changes and preparatory manipulation to emerge as task-dependent strategies rather than prescribed procedures. Our evaluation mainly focuses on three frontier multimodal agents that operate through a common visuomotor interface and are evaluated by physical task outcomes, interaction efficiency, and explicit task submission. We additionally provide full-episode trajectories to support learning-based methods and evaluate a trained policy , allowing zero-shot evaluation of out-of-the-box frontier robotic agents and trained vision-language-action policies under the same task framework. Our benchmark aims to study not only whether a robot can execute an action, but whether it can determine what it needs to learn from the environment, physically obtain that information, and use it to complete a goal. This also allows teasing out the failure points in the trajectory of task execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.